Search PubMedSearch

SEARCH · Search PubMed

Results for “Dark genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The dark genome in cardiovascular medicine.

Only ∼1%-2% of the human genome directly codes for proteins. The remainder consists of non-coding DNA, often referred to as the 'dark genome'. This includes regulatory elements, transposable and repetitive sequences, structural genomic features, pseudogenes, intronic and intergenic regions, and non-coding RNA (ncRNA) genes. These components are increasingly recognized as major regulators of gene expression, cell identity, and disease susceptibility. Currently, dark genome elements, particularly ncRNAs are increasingly recognized as important regulators of cardiovascular health and disease. Advances in genome analysis technologies have greatly improved our understanding of these non-coding regions and revealed clearer connections between the dark genome and cardiovascular traits. This review highlights major parts of the dark genome involved in cardiovascular disease, with emphasis on those for which mechanistic understanding and translational relevance are beginning to emerge. As mechanistic insight into individual and collective components of the dark genome advances, it increasingly enables the development of new opportunities for targeted therapeutics for cardiovascular prevention and disease management.

Humans

Machine learning approaches for cancer prognosis and diagnosis via non-coding RNA: a comprehensive review.

Non-coding RNAs (ncRNAs), once considered genomic dark matter, are now established as key regulators of gene expression with widespread roles in cellular homeostasis and disease. In cancer, ncRNA expression is frequently and systematically dysregulated, and many of these molecules circulate in stable, protected form within biofluids, offering a compelling basis for non-invasive or minimally invasive diagnostic strategies. However, their clinical translation remains substantially hindered to date due to biological complexity, technical noise, and high dimensionality inherent to ncRNA expression datasets. In this context, machine learning (ML) has emerged as a powerful analytical tool to address these challenges, enabling the identification of subtle, reproducible ncRNA signatures predictive of diverse malignancies. This review critically evaluates ML-driven frameworks for cancer diagnosis and prognosis across four ncRNA subclasses, namely miRNAs, lncRNAs, circRNAs, and piRNAs, while also acknowledging the biophysical and thermodynamic models that reinforce ncRNA bioinformatics. Despite substantial methodological progress in ML-based cancer diagnosis and prognosis, key challenges persist, including tumor biological heterogeneity, limited multicenter validation, and the lack of widely adopted standardized protocols for preprocessing, normalization, and reporting workflows. Furthermore, many current ML models lack interpretability in biological or clinical context, constraining their translational utility. By synthesizing recent advances and identifying unresolved barriers, this review charts a roadmap for developing a robust, clinically actionable ncRNA biomarker platform for cancer detection. With global cancer incidence projected to exceed 35 million annual cases by 2050, validated ncRNA-ML-driven frameworks hold potential to revolutionize early-stage detection and personalized therapeutic strategies, thereby reducing the escalating socio-economic burden of cancer worldwide.

Humans

Systematic discovery of retina-enriched Rik genes identifies 1190005I06Rik as a novel modulator of visual signalling.

BACKGROUND: High‑throughput transcriptome projects have revealed thousands of mammalian genes with little or no functional annotation. Among these are hundreds of loci assigned provisional “Rik” identifiers following discovery in the RIKEN cDNA annotation effort. Although often dismissed as genomic dark matter, such genes may encode tissue‑restricted proteins that modulate physiologic functions and influence disease. The retina is a highly specialised neural tissue and a common site of inherited disorders; understanding its molecular repertoire could illuminate novel therapeutic avenues. METHODS: We integrated bulk RNA‑seq from ten adult mouse tissues, evolutionary and domain analysis, single‑cell RNA‑seq, and CRISPR/Cas9 gene disruption to systematically catalogue protein‑coding Rik genes enriched in the retina and test the function of a representative gene. RESULTS: A rigorous differential expression analysis identified 44 Rik genes with robust retina‑specific expression compared with nine non‑retinal tissues. Many of these genes lack orthologues beyond rodents, while others show broad conservation, illustrating a continuum from lineage‑restricted to conserved retinopathy candidates. Single‑cell transcriptomics revealed that these genes are expressed across retinal cell types, with the highest aggregate expression in cone photoreceptors and inner interneurons. To evaluate physiological significance, we generated a 1190005I06Rik knockout mouse. Although retinal architecture appeared normal, loss of 1190005I06Rik enhanced electroretinogram b‑wave amplitudes and altered light‑avoidance behaviour, indicating that this previously uncharacterised gene acts as a negative modulator of visual signalling. CONCLUSIONS: We present a curated atlas of retina‑enriched Rik genes and demonstrate that 1190005I06RIK modulates retinal circuit function. This resource expands the molecular landscape of the retina and provides new candidates for the genetic basis of inherited retinal disease. Our findings underscore that unannotated genes may exert measurable effects on sensory processing and warrant systematic exploration in the context of human ocular disorders.

Animals

Nanobioreactor detection of space-associated hematopoietic stem and progenitor cell aging.

Human hematopoietic stem and progenitor cell (HSPC) fitness declines following exposure to stressors that reduce survival, dormancy, telomere maintenance, and self-renewal, thereby accelerating aging. While previous National Aeronautics and Space Administration (NASA) research revealed immune dysfunction in low-earth orbit (LEO), the impact of spaceflight on human HSPC aging had not been studied. To study HSPC aging, our NASA-supported Integrated Space Stem Cell Orbital Research (ISSCOR) team developed bone marrow niche nanobioreactors with lentiviral bicistronic fluorescent, ubiquitination-based cell-cycle indicator (FUCCI2BL) reporter for real-time HSPC tracking in artificial intelligence (AI)-driven CubeLabs. In month-long International Space Station (ISS) missions (SpX-24, SpX-25, SpX-26, and SpX-27) compared with ground controls, FUCCI2BL reporter, whole-genome and transcriptome sequencing, and cytokine arrays demonstrated cell-cycle, inflammatory cytokine, mitochondrial gene, human repetitive element, and apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like 3 (APOBEC3) deregulation together with clonal hematopoietic mutations. Furthermore, HSPC functionally organized multi-omics aging (HSPC-FOMA) analyses revealed reduced telomere maintenance, adenosine deaminase acting on RNA1 (ADAR1) p150 self-renewal gene expression, and replating capacity indicative of space-associated HSPC aging that may limit long-duration spaceflight.

Humans

Decoding the Functional Interactome of Non-Model Organisms with PHILHARMONIC.

Despite the widespread availability of genome sequencing pipelines, many genes remain part of the genome's "dark matter," where existing inference tools cannot even begin to guess the biological function of their proteins from sequence alone. This challenge is especially pronounced in organisms that are highly evolutionarily distant from well-studied models, where homology-based methods break down. Here, we describe PHILHARMONIC, a computational method that combines deep learning-based de novo protein interaction network inference with robust unsupervised spectral clustering and remote homology to illuminate functional organization in any non-model organism. From only a sequenced proteome, we show PHILHARMONIC predicts protein functions, functional communities, and higher-order network structure with high accuracy. We validate its performance using experimental gene expression and pathway data in D. melanogaster, and we demonstrate its broad utility by analyzing temperature sensing and stress response pathways in the reef-building coral P. damicornis and its algal symbiont C. goreaui. PHILHARMONIC provides a general-purpose engine for functional discovery and biological hypothesis generation in non-model organisms, enabling systems-level insights across the full diversity of life.

Journal Article

Characterization of microbial dark matter at scale with MetaSBT and taxonomy-aware Sequence Bloom Trees.

Metagenomics has become a powerful tool for studying microbial communities, allowing researchers to investigate microbial diversity within complex environmental samples. Recent advances in sequencing technology have enabled the recovery of near-complete microbial genomes directly from metagenomic samples, also known as metagenome-assembled genomes (MAGs). However, accurately characterizing these genomes remains a significant challenge due to the presence of sequencing errors, incomplete assembly, and contamination. Here we present MetaSBT, a new tool for organizing, indexing, and characterizing microbial reference genomes and MAGs. It is able to identify clusters of genomes at all seven taxonomic levels, from the kingdom all the way down to the species level, using the Sequence Bloom Tree (SBT) data structure that relies on Bloom Filters (BFs) to index massive amounts of genomes based on their k-mers composition. We have built an initial set of databases composed of over 190 thousand viral genomes from NCBI GenBank and public sources grouped into sequence consistent clusters at different taxonomic levels, making it the first software solution for the classification of viruses at different ranks, including still unknown ones. This results in the definition of over 40 thousand species clusters where ~80% do not match with any known viral species in reference databases to date. Furthermore, we show how our databases can be used as a new basis for existing quantitative metagenomic profilers to unlock the detection of unknown microbes and the estimation of their abundance in metagenomic samples. Finally, the framework is released open-source and, along with its public databases, is fully integrated into the Galaxy Platform enabling broad accessibility.

metagenome-assembled genomes

Genomic signatures of the Arctic-adapted North American gray wolf ecotype.

The Arctic Circle is one of Earth's most extreme environments. It features cold temperatures, resource shortages, and near-complete winter darkness. Here, we generated a chromosome-level genome assembly of a male wolf from the Arctic Archipelago (Canis lupus arctos). Our assembly and that of the related C. l. orion from Greenland was used to identify candidate genic and regulatory features of Arctic-adapted polar wolves, ranging from selection acting on standing variation and amino acid changes in genes to conserved non-exonic elements (CNEEs) that may regulate gene expression. We identified genes and nearby CNEEs associated with thermoregulation (e.g., APOB), coat color and patterning (e.g., GOLGB1), and DNA damage response (e.g., POLQ). In vitro assays supported changes in polar wolf gene (POLQ and TRPV2 amino acid substitutions) and CNEE function. Our report offers insights into the genetic mechanisms underlying polar wolf adaptations, laying a foundation for future studies on Arctic canines.

Amino acid substitutions

CoSAG-nf: A Scalable Nextflow Pipeline for Co-assembly, Optimization, and Interactive Visualization of High-Throughput Single-Cell Genomes.

MOTIVATION: Single-cell amplified genomes (SAGs) are crucial for resolving intra-population microbial heterogeneity and accurately understanding the metabolic potential of microbial dark matter populations. However, SAGs generated through multiple displacement amplification (MDA) of genomic DNA from single cells with single-copy chromosomes are highly fragmented and prone to contamination, severely hindering high-quality genome reconstruction and functional analysis, which greatly limits their scientific utility. Co-assembly of related SAGs can substantially improve genome quality, but to our knowledge no automated pipeline exists for high-throughput processing, forcing manual implementation of complex workflows that scale poorly to modern dataset sizes. RESULTS: We present CoSAG-nf, an automated high-throughput co-assembly and optimization pipeline for SAGs, implemented following the nf-core framework standards. The pipeline performs alignment-free clustering using sourmash MinHash signatures, then employs iterative tetranucleotide frequency profiling to identify and exclude outlier SAGs from co-assembly groups. CheckM2 quality assessment guides dynamic selection of optimal SAG combinations to optimize genome completeness and minimize contamination. Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms. The pipeline generates comprehensive HTML reports with quality metrics and taxonomic annotations, providing an end-to-end solution for automated high-throughput single-cell genome reconstruction. AVAILABILITY: CoSAG-nf is freely available under the MIT License at: https://github.com/linfengxu/CoSAG-nf. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.21525244. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Journal Article

Evidence of genome-wide relaxed selection on mildly deleterious mutations in an ancient subterranean catfish.

About one hundred subterranean catfish species have been described, resulting from repeated colonization of cave environments by multiple surface lineages. Most cave-dwelling species are found in the Americas, in particular in South America, but a few species also live in Central and North America. Despite the availability of high-quality genome assemblies for two cave species, the Mexican blind catfish Prietella phreatophila and the Colombian blind catfish Trichomycterus rosablanca, genomic approaches to investigate genetic changes associated with subterranean life or to estimate cave colonization times remain largely unexplored. To fill this gap, we additionally sequenced the genomes of four blind and depigmented subterranean catfishes from Peru (three Trichomycterus and one Astroblepus), as well as the genomes of four close surface relatives. We first extracted a large set of light-related genes, such as phototransduction and crystallin genes, and found contrasting decays of these sequences in different cave species, from 1% of pseudogenes in T. rosablanca to 48% in P. phreatophila. Two independent molecular dating methods gave congruent ages, indicating that these catfishes colonized subterranean habitats at different times, ranging from Early Pliocene to Late Pleistocene, supporting the hypothesis that surface catfishes repeatedly and rapidly adapted to subterranean habitats. The oldest cave species, P. phreatophila, appears to have been thriving in the dark for over 3.5 million years. Moreover, a genome-wide analysis of protein-coding genes suggests weaker purifying selection on mildly deleterious mutations in this cavefish than in other catfish lineages, likely reflecting a long-term small effective population size.

cavefishes

ECLIPSE: exploring the dark proteome of ESKAPE pathogens through the sequence similarity network of the Protein Universe Atlas.

MOTIVATION: The accelerating crisis of antimicrobial resistance among the critical so-called ESKAPE pathogens demands the urgent identification of novel molecular targets. However, a substantial fraction of ESKAPE proteomes remains functionally uncharacterized, with many genes annotated as encoding hypothetical proteins. These protein sequences often lack significant similarity to known protein families when conventional homology-based annotation methods are used and thus remain "dark". This limits our ability to explore their roles in pathogenicity, and it is thus crucial to bridge this substantial gap in pathogen biology by developing new strategies to illuminate these "dark" regions of the ESKAPE pan-proteome. RESULTS: We introduce ECLIPSE (ESKAPE Connectome Linkage and Inference for Proteome Sequence Exploration), a network-based computational framework that systematically identifies and prioritizes functionally dark protein families in ESKAPE pan-proteomes. ECLIPSE embeds target ESKAPE pathogen proteomes within the global sequence similarity network of the Protein Universe Atlas. It detects connected components composed entirely of unannotated proteins, called the "dark proteome." As a case study, we applied ECLIPSE to a pan-proteome of 3 460 657 protein sequences from 635 strains of Pseudomonas aeruginosa (PA). ECLIPSE identified 120 985 proteins (4%) residing in completely dark connected components. Furthermore, we have performed a taxonomic diversity analysis using normalized Shannon indices to characterize each dark component by its enrichment in ESKAPE pathogens. The analysis utilized the evenness (E) value (see Methods 2.1), which distinguishes Pseudomonas-specific (target-specific) from ESKAPE-enriched dark components. We then developed the Dark Proteome Prioritization Score (DPPS), a composite multidimensional scoring framework (see Methods 2.5). It ranks these dark components by biological relevance across four orthogonal axes: (i) functional darkness, (ii) P. aeruginosa proportion in the Atlas, (iii) AMR-clade taxonomic restriction, and (iv) conservation across the 635 P. aeruginosa strains. This framework outputs a robust four-tier scoring system; the prioritized Tier I components were validated by weight sensitivity analysis and remained stable across 500 Monte Carlo weight perturbations. Structural characterization of one of the top-ranked ESKAPE-enriched dark components revealed that it belongs to the beta-barrel fold DUF1302 (PF06980) family, for which no experimentally solved three-dimensional structure exists in the PDB. The genomic context analysis indicates that it is co-localized with a LuxR-type transcriptional regulator. Collectively, ECLIPSE identifies evolutionarily conserved, structurally defined, and functionally dark proteins enriched across ESKAPE pathogens; these dark proteins can further be utilized as alternative antimicrobial targets for experimental characterization. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available for free at: Github: https://github.com/surabhilata/ECLIPSE.git, Zenodo: DOI: 10.5281/zenodo.21064323.

Proteome

The mitochondrial genome of Chlamydomonas. II. Genetic analysis of non-mendelian obligate photautotrophic mutants.

Among a collection of obligate photoautotrophic (dark-dier, dk) mutants isolated in Chlamydomonas reinhardtii, two have been found which are inherited in crosses to wild type in a non-Mendelian, biparental and apparently random fashion. F1 progeny include not only cells which show the dk and wildtype parental phenotypes but also many which possess intermediate phenotypes between wild type and dk. When F1 progeny with dk, intermediate or wild-type phenotype were backcrossed to wild type, the dk phenotype continued to be inherited in a biparental and random fashion. Upon selection, neither mutant formed stable clones producing only dk progeny, suggesting that the two mutants segregate dk and wild-type progeny somatically and that the homozygous dk condition may be lethal. The biparental transmission of these two non-Mendelian dk mutations resembles the transmission of acriflavin-induced minute mutations of Chlamydomonas and is distinct from the uniparentally inherited chloroplast mutations of this alga. Both the dk and minute mutations may alter mitochondrial DNA and thereby alter mitochondrial functions.

Animals

Stepping out of the dark: how metabolomics shed light on fungal biology.

Metabolomics, a critical tool for analyzing small-molecule metabolites, integrates with genomics, transcriptomics, and proteomics to provide a systems-level understanding of fungal biology. By mapping metabolic networks, it elucidates regulatory mechanisms driving physiological and ecological adaptations. In fungal pathogenesis, metabolomics reveals host-pathogen dynamics, identifying virulence factors like gliotoxin in Aspergillus fumigatus and metabolic shifts, such as glyoxylate cycle upregulation in Candida albicans. Ecologically, it highlights fungal responses to abiotic stressors, including osmolyte production like trehalose, enhancing survival in extreme environments. These insights highlight metabolomics' role in decoding fungal persistence and niche colonization. In drug discovery, it aids target identification by profiling biosynthetic pathways, supporting novel antifungal and nanostructured therapy development. Combined with multi-omics, metabolomics advances insights into fungal pathogenesis, ecological interactions, and therapeutic innovation, offering translational potential for addressing antifungal resistance and improving treatment outcomes for fungal infections. Its progress shed light on complex fungal molecular profiles, advancing discovery and innovation in fungal biology.

Metabolomics

Repeated evolution on oceanic islands: comparative genomics reveals species-specific processes in birds.

Understanding the interplay between genetic drift, natural selection, gene flow, and demographic history in driving phenotypic and genomic differentiation of insular populations can help us gain insight into the speciation process. Comparing patterns across different insular taxa subjected to similar selective pressures upon colonizing oceanic islands provides the opportunity to study repeated evolution and identify shared patterns in their genomic landscapes of differentiation. We selected four species of passerine birds (Common Chaffinch Fringilla coelebs/canariensis, Red-billed Chough Pyrrhocorax pyrrhocorax, House Finch  Haemorhous mexicanus and Dark-eyed/island Junco Junco hyemalis/insularis) that have both mainland and insular populations. Changes in body size between island and mainland populations were consistent with the island rule. For each species, we sequenced whole genomes from mainland and insular individuals to infer their demographic history, characterize their genomic differentiation, and identify the factors shaping them. We estimated the relative (Fst) and absolute (dxy) differentiation, nucleotide diversity (π), Tajima's D, gene density and recombination rate. We also searched for selective sweeps and chromosomal inversions along the genome. All species shared a marked reduction in effective population size (Ne) upon island colonization. We found diverse patterns of differentiated genomic regions relative to the genome average in all four species, suggesting the role of selection in island-mainland differentiation, yet the lack of congruence in the location of these regions indicates that each species evolved differently in insular environments. Our results suggest that the genomic mechanisms involved in the divergence upon island colonization-such as chromosomal inversions, and historical factors like recurrent selection-differ in each species, despite the highly conserved structure of avian genomes and the similar selective factors involved. These differences are likely influenced by factors such as genetic drift, the polygenic nature of fitness traits and the action of case-specific selective pressures.

Animals

Auditing bacterial dark-gene screens for superimposed open reading frame artefacts: A multi-layer analysis of Rv2438A in Mycobacterium tuberculosis.

Essentiality and knockdown-vulnerability screens can promote spurious bacterial open reading frames when those frames overlap essential genes, because such a frame inherits its neighbour's signals undiluted and therefore satisfies the screen's criteria better than a genuine small gene. We present a multi-layer audit that tests this failure mode across genome annotation, transposon mutagenesis, CRISPR interference, homology, transcript mapping, proteomics, and population variation. We apply it to Rv2438A, a 92-codon conserved hypothetical open reading frame of Mycobacterium tuberculosis ranked first by our own dark-gene target screen. Rv2438A is superimposed on the essential NAD synthetase locus nadE: 44% lies within its coding sequence on the opposite strand, and the remainder covers its promoter and transcription start site. Consequently, three of five Himar1 sites lie within nadE, no CRISPRi guide can target Rv2438A without binding nadE, and the cross-species hit maps to the same nadE start junction. Rv2438A lacks its own transcription start site and is absent from every proteomic dataset that detects nadE. A genome-wide scan identifies six short, overlapping, uncharacterised loci among 3907 annotated genes, but only Rv2438A combines overlap and essentiality with non-detection across all proteomic datasets; rare genome-wide, it ranked first among screen hits. We provide an implementable audit workflow and a codon-position control, but measure the control's sensitivity as only two of five genes with attested protein, limiting it to confirmatory use. Overlap coordinates and neighbour-specific experimental resolution should therefore be reported before bacterial dark genes are prioritised.

CRISPR interference

Seed2LP: seed inference in metabolic networks for reverse ecology applications.

MOTIVATION: A challenging problem in microbiology is to determine nutritional requirements of microorganisms and culture them, especially for the microbial dark matter detected solely with culture-independent methods. The latter foster an increasing amount of genomic sequences that can be explored with reverse ecology approaches to raise hypotheses on the corresponding populations. Building upon genome-scale metabolic networks (GSMNs) obtained from genome annotations, metabolic models predict contextualized phenotypes using nutrient information. RESULTS: We developed the tool Seed2LP, addressing the inverse problem of predicting source nutrients, or seeds, from a GSMN and a metabolic objective. The originality of Seed2LP is its hybrid model, combining a scalable and discrete Boolean approximation of metabolic activity, with the numerically accurate flux balance analysis (FBA). Seed inference is highly customizable, with multiple search and solving modes, exploring the search space of external and internal metabolites combinations. Application to a benchmark of 107 curated GSMNs highlights the usefulness of a logic modelling method over a graph-based approach to predict seeds, and the relevance of hybrid solving to satisfy FBA constraints. Focusing on the dependency between metabolism and environment, Seed2LP is a computational support contributing to address the multifactorial challenge of culturing possibly uncultured microorganisms. AVAILABILITY AND IMPLEMENTATION: Seed2LP is available on https://github.com/bioasp/seed2lp.

Metabolic Networks and Pathways

There is gold in the graveyard: a new lineage of zombie-ant fungi in the genus Ophiocordyceps (Ophiocordycipitaceae: Hypocreales) from Minas Gerais, Brazil.

Ophiocordyceps serves as a key model for studying cryptic fungal diversity and behavioural manipulation of hymenopterous insects. Here, we describe Ophiocordyceps acanthoponerae, a newly discovered species infecting Acanthoponera mucronata (Heteroponerini: Formicidae) in a Brazilian Atlantic rainforest-Cerrado ecotone. Morphological analyses revealed mixed traits characteristic of Ophiocordyceps lineages associated with ants and wasps, including leaf biting behaviour manipulation, dark brown ascostromata covering 360º of the stalk, ascospores producing capilliconidia and hirsutelloid asexual morphs. Phylogenetic analyses based on four genomic regions (SSU, LSU, TEF and RPB1) placed this species outside the traditional myrmecophilous hirsutelloid clades O. unilateralis and O. kniphofioides, and within a novel clade closely related to the wasp pathogen O. humbertii. This discovery represents the first record of Ophiocordyceps infecting Heteroponerini and highlights an unexplored lineage of manipulative fungi. Our findings expand the known host range for myrmecophilous Ophiocordyceps and underscore the importance of studying fungal diversity in under-sampled ecological niches. Citation: Lima-Santos SJ, Araújo JPM, Feitosa RM, Mendes-Pereira T, Elliot SL, Evans HC (2025). There is gold in the graveyard: a new lineage of zombie-ant fungi in the genus Ophiocordyceps (Ophiocordycipitaceae: Hypocreales) from Minas Gerais, Brazil. Fungal Systematics and Evolution 16: 243-264. doi: 10.3114/fuse.2025.16.14.

Acanthoponera

Doubled Genomes, Divergent Fates: Genomic Insights Into Diversification in an Allotetraploid Cavefish.

Cave environments impose unique challenges that drive remarkable genetic and phenotypic changes in cave-dwelling organisms. In this study, we investigated the genomic basis of adaptation in the small eye golden-line fish (Sinocyclocheilus microphthalmus), an allotetraploid cavefish endemic to Guangxi, China. Using whole-genome resequencing data from 47 individuals across six cave locations, we examined how neutral and selective forces influence diversification. Our analyses uncovered significant population structure indicative of allopatric divergence, along with evidence of locus-specific selection contributing to genomic differentiation. We identified seven single outlier clusters (SOCs), each tied to the divergence of specific populations, underscoring the role of local processes in driving diversity. Genes associated with vision showed relaxed selection, likely reflecting adaptation to darkness, while positive selection on other loci revealed additional functional shifts. Notably, allopolyploidy was found to fuel divergence through subgenome-specific patterns and asymmetric evolution within SOCs and among homoeologs. Taken together, these findings provide valuable insights into mechanisms of cave evolution and illustrate how allotetraploid genomes can facilitate diversification, potentially contributing to speciation in extreme environments.

Animals