Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

The genome sequence of Quercus petraea Liebl., 1784 (Fagales: Fagaceae).

We present a genome assembly of Quercus petraea (sessile oak; Streptophyta; Magnoliopsida; Fagales; Fagaceae). The assembly consists of two haplotypes with total lengths of 823.57 megabases and 815.09 megabases. Most of haplotype 1 (99.41%) is scaffolded into 12 chromosomal pseudomolecules. Haplotype 2 was assembled to scaffold level. The mitochondrial sequence has a length of 349.29 kilobases and the plastid genome assembly has a length of 161.18 kilobases. Gene annotation of this assembly on Ensembl identified 32 253 protein-coding genes. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Fagales↗

Chromosome-level genome assembly of hawthorn spider mite, Amphitetranychus viennensis (Acari: Tetranychidae).

The hawthorn spider mite, Amphitetranychus viennensis, is a major pest of orchards and ornamentals in the Palaearctic region, with adaptability and acaricide resistance. The lack of high-quality genomic resources limits understanding of its detoxification mechanisms and the development of RNAi-based pest control strategies. In this study, we utilized Illumina, Pacific Biosciences (PacBio), and Hi-C sequencing technologies to assemble a chromosome-level reference genome of A. viennensis. The assembled genome spans 141.96 Mb, with a contig N50 of 1.35 Mb. BUSCO analysis confirmed a high level of completeness, covering 91.6% of annotated genes. The assembly includes 50.97 Mb of repetitive sequences, representing 35.93% of the genome, and annotates 13,968 protein-coding genes. Using Hi-C sequencing, we anchored 47 contigs to three chromosomes, accounting for 97.27% of the estimated nuclear genome and achieving a contig N50 of 45.83 Mb. This high-quality genome assembly provides a valuable foundation for evolutionary and genomic research on spider mites, while also serving as a genetic resource to inform molecular control strategies and support sustainable pest management.

Animals↗

Genome-based predictions of metabolic preferences and substrate phenotypes in psychrotrophic bacteria from permafrost environments.

Genomes reveal vast functional potential, but harbor genomic noise that obscures prediction of metabolic and environmental preferences. Genomic databases are skewed towards clinically relevant and easily cultivated bacteria, limiting predictions for diverse and underrepresented environmental taxa. Psychrotrophic bacteria, which can survive and grow in cold, nutrient-limited, dry, and saline environments, are especially underrepresented despite their relevance for understanding microbial responses to changing cold environments and potential biotechnological value given growth at low temperatures. Assembling complete genomes of 48 isolates from Alaskan permafrost, seasonally frozen active layer soils, and terrestrial ice, we used Kyoto Encyclopedia of Genes and Genomes (KEGG) ortholog annotations to evaluate the predictability of metabolic resource-use traits observed using phenotypic tests. Genome-predicted values for glycolytic versus gluconeogenic catabolic preference index, or sugar-acid preference (SAP), explained over 50% of the variance in empirically observed SAP. SAP was inversely correlated to genomic GC content, which follows phylum-level trends, indicating that coarse metabolic preference covaries with phylogeny. Regularized elastic net models offered a more granular view, linking KEGG genes to specific substrate utilization and sensitivity phenotypes and yielding moderate but reproducible accuracy (AUC 0.70-0.79) for 11 substrates, demonstrating that specific substrate responses may be predictable from relatively small subsets of KO genes. These results extend recent advances, such as the SAP metric, and highlight associations among genomic GC content, phylum, and broad metabolic strategy. Linking genomic content to phenotype using isolates is a necessary step toward predictive models of microbial function in environmental communities, and this work can be used for hypothesis generation, with applications towards more expansive data sets.IMPORTANCECold region soils and ice host psychrotrophic bacteria with metabolic traits and adaptations that enable persistence in harsh, resource-limited environments. However, these taxa are underrepresented in genomic reference databases dominated by well-studied, mesophilic organisms. This gap limits inference of ecological strategies and our ability to predict how these microbes may influence the large, thaw-vulnerable carbon reservoirs in permafrost. Here, we show that genomic GC content is associated with the sugar-versus-acid catabolic preference (SAP) of isolates across major phyla, suggesting that broad genomic features may provide a coarse signal of metabolic strategy. We demonstrate that a modified SAP metric, using binary (positive/negative) substrate utilization rather than detailed growth rate measurements, is moderately predictive, thus extending its application to slow-growing or difficult-to-culture taxa. Together, these advances broaden the toolkit for linking genome content to resource-use traits (phenotype) in poorly characterized, cold-adapted bacteria and offer a tractable entry point to broad prediction and hypothesis generation.

Genome, Bacterial↗

The genome sequence of the bronze furrow bee, Seladonia tumulorum (Linnaeus, 1758).

We present the haploid genome assembly of an individual male Seladonia tumulorum (the bronze furrow bee; Arthropoda; Insecta; Hymenoptera; Halictidae). The genome sequence is 479 megabases in span. Most of the assembly (84.28%) is scaffolded into 17 chromosomal pseudomolecules. The mitochondrial genome was also assembled and is 17.3 kilobases in length. Gene annotation of this assembly on Ensembl identified 19,308 protein-coding genes. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Hymenoptera↗

Using intrahost single nucleotide variant data to predict SARS-CoV-2 detection cycle threshold values.

Over the last four years, each successive wave of the COVID-19 pandemic has been caused by variants with mutations that improve the transmissibility of the virus. Despite this, we still lack tools for predicting clinically important features of the virus. In this study, we show that it is possible to predict the PCR cycle threshold (Ct) values from clinical detection assays using sequence data. Ct values often correspond with patient viral load and the epidemiological trajectory of the pandemic. Using a collection of 36,335 high quality genomes, we built models from SARS-CoV-2 intrahost single nucleotide variant (iSNV) data, computing XGBoost models from the frequencies of A, T, G, C, insertions, and deletions at each position relative to the Wuhan-Hu-1 reference genome. Our best model had an R2 of 0.604 [0.593-0.616, 95% confidence interval] and a Root Mean Square Error (RMSE) of 5.247 [5.156-5.337], demonstrating modest predictive power. Overall, we show that the results are stable relative to an external holdout set of genomes selected from SRA and are robust to patient status and the detection instruments that were used. This study highlights the importance of developing modeling strategies that can be applied to publicly available genome sequence data for use in disease prevention and control.

SARS-CoV-2↗

Detection of House Dust Mite-derived DNA in Human Lung Tumors by Whole-Genome Sequencing.

Lung cancer in never-smokers (LCINS) accounts for an increasing proportion of lung cancer cases, yet its risk factors remain poorly understood. House dust mites (HDM) are common aeroallergens that induce airway inflammation, but their potential contribution to lung cancer is unknown. We analyzed unmapped whole-genome sequencing reads from 783 lung cancers from the Sherlock-Lung (n = 621 never-smokers) and EAGLE (n = 162 smokers) cohorts, including 328 matched adjacent normal lung tissues. After removal of human sequences, reads were aligned to reference genomes from the two major HDM species and confirmed by BLAST. Samples with top BLAST matches were classified as HDM-detected. Associations between HDM detection and genomic, microbiome, and bulk RNA-seq-derived immune features were evaluated. HDM-derived DNA was detected at low abundance in a subset of tumors and adjacent normal tissues, with higher detection frequencies in tumors than matched normal tissues and in smokers than never-smokers. In LCINS tumors, HDM detection was not associated with tumor mutational burden or recurrent driver alterations but was associated with modest differences in immune cell composition and a limited but reproducible bacterial co-detection pattern. These findings provide a foundation for investigating aeroallergen-derived DNA signatures and their potential relationship to the lung tumor microenvironment.

Environmental exposure↗

Mapping multiple co-sequenced T-DNA integration sites within the Arabidopsis genome.

MOTIVATION: Insertion mutagenesis, using transgenes or endogenous transposons, is a popular method for generating null mutations (knockouts) in model organisms. Insertions are mapped to specific genes by amplifying (via TAIL-PCR) and sequencing genomic regions flanking the inserted DNA. The presence of multiple TAIL-PCR templates in one sequencing reaction results in chimeric sequence of intermittently low quality. Standard processing of this sequence by applying Phred quality requirements results in loss of informative sequence, whereas not trimming low-quality sequence causes inclusion of low-complexity homopolymers from the ends of sequence runs. Accurate mapping of the flanking sequences is complicated by the presence of gene families. RESULTS: Methods for extracting informative regions from sequence traces obtained by sequencing multiple TAIL-PCR fragments in a single reaction are described. The completely sequenced Arabidopsis genome was used to identify informative TAIL-PCR sequence regions. Methods were devised to define and select high quality matches and precisely map each insert to the correct genome location. These methods were used to analyze sequence of TAIL-PCR-amplified flanking regions of the inserts from individual plants in a T-DNA-mutagenized population of Arabidopsis thaliana, and are applicable to similar situations where a reference genome can be used to extract information from poor-quality sequence.

Arabidopsis↗

The genome sequence of an ichneumonid wasp, Venturia canescens (Gravenhorst, 1829) (Hymenoptera: Ichneumonidae).

We present a genome assembly from an individual female Venturia canescens (ichneumonid wasp; Arthropoda; Insecta; Hymenoptera; Ichneumonidae). The genome sequence has a total length of 299.96 megabases. Most of the assembly (97.25%) is scaffolded into 11 chromosomal pseudomolecules. The mitochondrial genome has also been assembled, with a length of 27.35 kilobases. Gene annotation of this assembly on Ensembl identified 14 281 protein-coding genes. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Hymenoptera↗

Comparative studies of the Acinetobacter genus and the species identification method based on the recA sequences.

The recA gene is indispensable for a maintaining and diversification of the bacterial genetic material. Given its important role in ensuring cell viability, it is not surprising that the RecA protein is both ubiquitous and well conserved among a range of prokaryotes. Previously, we reported Acinetobacter genomic species identification method based on PCR amplification of an internal fragment of the recA gene with subsequent restriction analysis (RFLP) with HinfI and MboI enzymes. In present study, the PCR products containing the internal fragment of the recA gene, for 25 Acinetobacter strains belonging to all genomic species, were sequenced. Based on the nucleotide sequences the restriction maps and phylogenetic tree were prepared. The restriction maps revealed that Tsp509I restriction enzyme is the most discriminating for RFLP. To verify the computer analysis, the amplified DNAs from all reference genomic species available (43 strains) and 34 clinical strains were digested with each of the three restriction endonucleases mentioned. The results of digestion confirmed the computer analysis. The reconstructed phylogenetic tree showed linkages between genomic species 1 (Acinetobacter calcoaceticus), 2 (Acinetobacter baumannii), 3, 'between 1 and 3', TU13 and 'close to TU13'; genomic species 4, 6, BJ13, BJ14, BJ15, BJ16 and BJ17; genomic species 7 (Acinetobacter johnsonii) and TU14; genomic species 10 and 11; genomic species 8 (Acinetobacter Iwoffii), 9, 12 (Acinetobacter radioresistens) and TU15; and genomic species 5 (Acinetobacter junii). It is interesting that one branch in the phylogenetic tree contains haemolytic species-genomic species 4 (A. haemolyticus), BJ13, BJ14, BJ15, BJ16 and BJ17. The proposed genotypic method clearly revealed that the RFLP profiles obtained with Tsp509I enzyme might be useful for species identification of Acinetobacter strains. In this context, recA/RFLP genotypic method should be seen as an ideal preliminary screening method for large numbers of isolates, with the ultimate confirmatory role reserved for DNA hybridization analysis.

Acinetobacter↗

Chromosome-level genome assembly of starry flounder (Platichthys stellatus).

Starry flounder (Platichthys stellatus) is widely distributed along the coastlines of the North Pacific. As an euryhaline flatfish, it can adapt to a wide range of environmental salinity ranging from freshwater to seawater, and is a promising aquaculture flatfish species in Korea and North China. However, no high-quality starry flounder reference genome has been reported to date, which greatly limits the studies of genetics and functional genomics. Here, we obtained a high-quality chromosome-level starry flounder genome assembly with a length of 643.56 Mb (scaffold N50: 26.19 Mb, contig N50: 10.00 Mb) combining short-reads sequencing, PacBio HiFi sequencing, and Hi-C sequencing. Approximately 94.02% of assembled sequences were anchored into 24 pseudochromosomes, and a total of 18 telomeres were detected. Totally 22,835 protein-coding genes and 227.87 Mb repetitive sequences were identified. In summary, the high-quality chromosome-level genome assembly not only provides valuable resources for genetic research in starry flounder, but also advances the development of molecular breeding technology of starry flounder.

Animals↗

Stability of HIV type 1 proviral genomes that contain two distinct primer-binding sites.

The initiation of human immunodeficiency virus type 1 (HIV-1) reverse transcription occurs by the extension of a tRNALys,3 positioned at an 18-nucleotide sequence in the RNA genome referred to as the primer-binding site (PBS). We have found that mutations within the PBS and a region upstream in U5, designated the A loop, influenced the selection of the tRNA primer used to initiate reverse transcription. Surprisingly, a proviral genome that contained a PBS and A loop complementary to tRNAPro resulted in the generation of viruses that contained two PBSs within the same genome: one of the PBSs in the virus was complementary to tRNALys,3 while the second PBS was complementary to tRNAIle, tRNAPro, or tRNALys,3. There were 14 nucleotides separating the two PBSs in the viral genome. In the current study, DNA encompassing U5 and the dual PBS complementary to the different tRNAs were amplified by PCR and exchanged for the corresponding region in an infectious HIV-1 clone, HXB2. Transfection of the different proviruses into cells resulted in the production of viruses that were infectious as determined by coculture with SupT1 cells. PCR was used to amplify the PBS regions from the different proviral DNAs followed by DNA sequencing of individual PCR clones. Proviruses containing the dual PBS complementary to tRNALys,3 and tRNAIle stably maintained the dual PBS complementary to both of these tRNAs following in vitro culture, although we noted consistent G-to-T and AA-to-GG substitutions in the 14-nucleotide region between the PBSs. The viruses derived from genomes that contained the dual PBS complementary to tRNALys,3 and tRNAPro also maintained both PBSs following in vitro culture; a single mutation was noted after in vitro culture in the 14-nucleotide region between the PBSs, which changed a consensus integration site (CA dinucleotide) prior to the PBS complementary to tRNAPro. In contrast, the proviral genomes containing the dual PBS complementary to tRNALys,3 were not stable and reverted back to a single PBS complementary to tRNALys,3. The results of our studies suggest that only the 5'-proximal PBS has been used to initiate reverse transcription. On the basis of our results, a mechanism is proposed for the generation of a dual PBS, which provides new insights into HIV-1 reverse transcription.

Animals↗

Whole Genome Sequencing Reveals How Plasticity and Genetic Differentiation Underlie Sympatric Morphs of Arctic Charr.

Salmonids have a remarkable ability to form sympatric morphs after postglacial colonisation of freshwater lakes. These morphs often differ in morphology, feeding and spawning behaviour. Here, we explored the genetic basis of morph differentiation in Arctic charr (n = 283) by first establishing a high-quality reference genome and then using this in whole genome sequencing of distinct morphs present in two Norwegian and two Icelandic lakes. The four lakes represent the spectrum of genetic differentiation between morphs from one lake with no genetic differentiation between morphs, implying phenotypic plasticity, to two lakes with locus-specific genetic differentiation, implying incomplete reproductive isolation, and one lake with strong genome-wide divergence consistent with complete reproductive isolation. As many as 12 putative inversions ranging from 0.45 to 3.25 Mbp in size segregated among the four morphs present in one lake, Thingvallavatn, and these contributed significantly to the genetic differentiation among morphs. None of the putative inversions were found in any of the other lakes, but there were cases of partial haplotype sharing in similar morph contrasts in other lakes. Our findings are consistent with a highly polygenic basis of morph differentiation with population-specific selection on alleles linked to the development of similar morph phenotypes. The results support a model where morph differentiation is first established through phenotypic plasticity, leading to niche expansion and separation. This may be followed by gradual development of reproductive isolation, locus-specific differentiation and eventually complete reproductive isolation and genome-wide divergence.

Whole Genome Sequencing↗

Refined phylogenetic profiles method for predicting protein-protein interactions.

MOTIVATION: The increasing availability of complete genome sequences provides excellent opportunity for the further development of tools for functional studies in proteomics. Several experimental approaches and in silico algorithms have been developed to cluster proteins into networks of biological significance that may provide new biological insights, especially into understanding the functions of many uncharacterized proteins. Among these methods, the phylogenetic profiles method has been widely used to predict protein-protein interactions. It involves the selection of reference organisms and identification of homologous proteins. Up to now, no published report has systematically studied the effects of the reference genome selection and the identification of homologous proteins upon the accuracy of this method. RESULTS: In this study, we optimized the phylogenetic profiles method by integrating phylogenetic relationships among reference organisms and sequence homology information to improve prediction accuracy. Our results revealed that the selection of the reference organisms set and the criteria for homology identification significantly are two critical factors for the prediction accuracy of this method. Our refined phylogenetic profiles method shows greater performance and potentially provides more reliable functional linkages compared with previous methods.

Algorithms↗

[Transcriptomes for serial analysis of gene expression].

The availability of the sequences for whole genomes is changing our understanding of cell biology. Functional genomics refers to the comprehensive analysis, at the protein level (proteome) and at the mRNA level (transcriptome) of all events associated with the expression of whole sets of genes. New methods have been developed for transcriptome analysis. Serial Analysis of Gene Expression (SAGE) is based on the massive sequential analysis of short cDNA sequence tags. Each tag is derived from a defined position within a transcript. Its size (14 bp) is sufficient to identify the corresponding gene and the number of times each tag is observed provides an accurate measurement of its expression level. Since tag populations can be widely amplified without altering their relative proportions, SAGE may be performed with minute amounts of biological extract. Dealing with the mass of data generated by SAGE necessitates computer analysis. A software is required to automatically detect and count tags from sequence files. Criterias allowing to assess the quality of experimental data can be included at this stage. To identify the corresponding genes, a database is created registering all virtual tags susceptible to be observed, based on the present status of the genome knowledge. By using currently available database functions, it is easy to match experimental and virtual tags, thus generating a new database registering identified tags, together with their expression levels. As an open system, SAGE is able to reveal new, yet unknown, transcripts. Their identification will become increasingly easier with the progress of genome annotation. However, their direct characterization can be attempted, since tag information may be sufficient to design primers allowing to extend unknown sequences. A major advantage of SAGE is that, by measuring expression levels without reference to an arbitrary standard, data are definitively acquired and cumulative. All publicly available data can thus be stored in a unique database, facilitating whole-genome analysis of differential expression between cell types, normal and diseased samples, or samples with and without drug treatment. SAGE data are readily amenable to statistical comparisons, allowing to determine the level of confidence of the observed variations. A major limitation of SAGE is that, because each analysis is obligatory performed on the whole set of expressed genes, it can hardly be performed on multiple samples, for example in kinetics studies or to compare the effects of large numbers of drugs. To overcome this limitation, high-throughput detection of a subset of mRNAs is more rapidly performed by parallel hybridization of mRNAs on arrays of nucleic acids immobilized on solid supports. From this point of view, a SAGE platform is a powerful instrument for selecting the most informative subset of genes, assembling them to design microarrays dedicated to a specific problem and calibrating measurement by comparison with a standard cell model for which SAGE data are available. This approach is an attractive alternative to strategies based exclusively on pangenomic arrays. A very large amount of SAGE data are already available and the problem is now to extract their biological meaning. Knowledge on metabolic pathways is already organized so that its successful integration in a SAGE platform can be undertaken. For other cell components and pathways, the problem lies on the lack of controlled vocabulary to describe gene activities, starting form a clear definition of the concept of biological function itself. Progress in gene and cell ontology is expected to facilitate computer-based extraction of biological knowledge from existing and forthcoming SAGE data.

Animals↗

ShiBASE: an integrated database for comparative genomics of Shigella.

Among the major enteric bacterial pathogens, Shigella is found to display extreme genome diversity and dynamics, which imposes a challenge in comparative genomic studies. To facilitate further studies in this area, we have constructed an integrated online database, ShiBASE (http://www.mgc.ac.cn/ShiBASE/),which contains Shigella genomic sequences of four species and additional comparative genomic hybridization (CGH) data of 43 serotypes. ShiBASE offers online comparative analysis on DNA sequences, gene orders, metabolic pathways and virulence factors. In addition, ShiBASE has a newly developed online comparative visualization service, Shi-align, which enables the alignment of any query sequence with the reference genome sequences.

Databases, Nucleic Acid↗

Direct RNA nanopore sequencing of full-length coronavirus genomes provides novel insights into structural variants and enables modification analysis.

Sequence analyses of RNA virus genomes remain challenging owing to the exceptional genetic plasticity of these viruses. Because of high mutation and recombination rates, genome replication by viral RNA-dependent RNA polymerases leads to populations of closely related viruses, so-called "quasispecies." Standard (short-read) sequencing technologies are ill-suited to reconstruct large numbers of full-length haplotypes of (1) RNA virus genomes and (2) subgenome-length (sg) RNAs composed of noncontiguous genome regions. Here, we used a full-length, direct RNA sequencing (DRS) approach based on nanopores to characterize viral RNAs produced in cells infected with a human coronavirus. By using DRS, we were able to map the longest (∼26-kb) contiguous read to the viral reference genome. By combining Illumina and Oxford Nanopore sequencing, we reconstructed a highly accurate consensus sequence of the human coronavirus (HCoV)-229E genome (27.3 kb). Furthermore, by using long reads that did not require an assembly step, we were able to identify, in infected cells, diverse and novel HCoV-229E sg RNAs that remain to be characterized. Also, the DRS approach, which circumvents reverse transcription and amplification of RNA, allowed us to detect methylation sites in viral RNAs. Our work paves the way for haplotype-based analyses of viral quasispecies by showing the feasibility of intra-sample haplotype separation. Even though several technical challenges remain to be addressed to exploit the potential of the nanopore technology fully, our work illustrates that DRS may significantly advance genomic studies of complex virus populations, including predictions on long-range interactions in individual full-length viral RNA haplotypes.

Cell Line↗

Genome-wide identification, structural characterization, and evolutionary analysis of growth-related gene families in African catfish (Clarias gariepinus).

The somatotropic axis encompassing growth hormone (GH), insulin-like growth factor (IGF), myostatin (MSTN), and prolactin (PRL) signalling cascades is the master regulator of somatic growth, metabolism, and development in vertebrates. African catfish (Clarias gariepinus), a commercially pivotal aquaculture species, now possesses a chromosome-level reference genome (CGAR_prim_01v2); however, a systematic, genome-wide characterization spanning all five interconnected growth-related gene families has not previously been undertaken in this species. Here, we identified and characterized 15 growth-related genes spanning gh1, ghra, ghrb, Igf1, Igf2a, Igf2b, igf1ra, Igf1rb, Igf2r, Mstna, Mstnb, prl, prlra, prlrb, and smtlb distributed across 13 chromosomes. Complete one-to-one orthology with zebrafish confirmed strong dosage-balance conservation across >120 million years of teleost divergence. Physicochemical analysis resolved a clear biochemical dichotomy between compact, basic secreted ligands (19.88-45.81 kDa; pI up to 10.02) and large, acidic, heavily glycosylated membrane receptors (56.82-270.80 kDa; pI 4.85-5.97). Phylogenetic analysis confirmed 3R whole-genome duplication origins for all paralog pairs, while synteny analysis revealed a disruption of the ancestral gh1-prl chromosomal block in C. gariepinus, a finding that warrants further comparative and functional investigation. This genomic atlas provides the sequence and structural information including exon-intron boundaries, domain architecture, and chromosomal coordinates needed as a prerequisite for future marker-assisted selection and CRISPR-based myostatin-editing efforts in African catfish aquaculture, though translation into applied breeding outcomes will require subsequent functional and expression studies.

Animals↗

Hidden genomic structure and widespread structural polymorphism across environmental gradients in the spiny sea star Marthasterias glacialis.

Genomic regions of reduced recombination can preserve linkage among co-adapted alleles, facilitating local adaptation despite high connectivity. Such regions-often generated by chromosomal inversions-may be especially important in highly dispersive marine taxa yet remain poorly documented in echinoderms. Here, we combined a chromosome-level reference genome with genome-wide ddRAD-seq from 296 Marthasterias glacialis individuals across 19 Atlantic-Mediterranean locations to quantify population structure and scan for recombination-suppressed haploblocks. Genome-wide neutral markers showed significant population differentiation together with evidence of high connectivity, revealed by the presence of inter-ecoregion migrants. Additionally, we identified 16 polymorphic haploblocks with patterns consistent with putative chromosomal inversions spanning 18.6% of the genome. Haploblock haplotypes were strongly environmentally and geographically structured and contained genes with key functions in stress response, osmoregulation and thermal tolerance. Haplotype distributions also paralleled previously described mitochondrial lineages despite nuclear gene flow, consistent with a model of ancient divergence followed by secondary contact. Overall, our results suggest a role for widespread structural polymorphism in adaptive differentiation in Echinodermata, providing a framework for linking echinoderm genome rearrangements to ecological divergence. Marthasterias glacialis thus emerges as a promising system to explore how structural variation contributes to adaptation and genome evolution in highly dispersive organisms.

Animals↗