Search PubMedSearch

SEARCH · Search PubMed

Results for “Genome assembly”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Chromosome-scale genome assembly and genomic prediction of essential oil compounds in Atractylodes lancea for genomics-assisted breeding.

Atractylodes lancea rhizomes are used as crude drugs. Essential oil compounds, including atractylodin, hinesol, β-eudesmol, and atractylon, are key determinants of crude drug quality. Conventional breeding of A. lancea is difficult because of its perennial growth. In this study, a chromosome-scale reference genome of A. lancea (4.79 Gb) was generated, and genome-wide association studies (GWAS) and genomic predictions of essential oil compounds were conducted to explore the potential for genome-assisted breeding. Genotyping of 480 lines using double-digest restriction-site-associated DNA-sequencing yielded 29,136 high-quality SNPs. All the compounds showed high genomic heritability (h2 = 0.758-0.915), indicating strong genetic control. Despite the high genomic heritability, GWAS detected only one weak association with atractylon and no significant loci for the three compounds. However, genomic prediction achieved moderate to high accuracy across multiple models, particularly the ridge regression, genomic best linear unbiased prediction, and Bayesian approaches. The prediction accuracy, measured as the Pearson correlation coefficient between the observed and predicted values, exceeded 0.6 for all four essential oil compounds. These results demonstrate the efficacy of genomic selection for improving essential oil compound levels in A. lancea and provide a foundation for genome-assisted breeding of medicinal plants with long breeding cycles.

Atractylodes lancea

Reference-Guided Chromosome-Scale Genome Assembly With Insights on Population Genomics of the Atlantic Goliath Grouper (Epinephelus itajara), Islas del Rosario, Colombia.

Epinephelus itajara, commonly known as the Atlantic Goliath grouper, is the largest species among the western North Atlantic groupers and is critically endangered. This species plays a crucial ecological, cultural, and economic role and has been the focus of captive breeding efforts at the Oceanario of the Rosario Islands, Colombia. However, despite its ecological and conservation importance, genomic resources and population genomic data for E. itajara remain scarce, particularly in the Colombian Caribbean. This study presents a reference-guided chromosome-scale genome assembly and an analysis of the population genomic structure of E. itajara using PacBio HiFi sequencing and Illumina technologies. The assembled genome has a total size of 1.12 Gb, with a contig N50 of 42.69 Mb and a scaffold N50 of 46.30 Mb. A total of 22,692 protein-coding genes were identified after masking 46% of the genome, which consists of repetitive elements. Comparative genomic analyses revealed a high degree of collinearity with closely related Epinephelus species and identified E. lanceolatus as the closest relative, supporting recent divergence and conserved genome architecture within the genus. Additionally, a population genomics analysis was conducted using 7706 high-quality SNPs to assess the genomic structure of captive populations. The results revealed four distinct genomic lineages, with moderate genetic differentiation among the sampled individuals. In the Colombian Caribbean, two unique lineages were identified, associated with the localities of Bahía Cispatá and Bahía Barbacoas, suggesting possible geographic isolation. These genomic resources provide valuable tools and new opportunities to better understand the genomic diversity, evolutionary history, and reproductive mechanisms of E. itajara. Moreover, they serve as a foundation for conservation strategies, including selective breeding programs aimed at increasing genomic diversity in captive populations and guiding restoration efforts in its natural habitat.

Epinephelus itajara

Chromosome-level genome assembly of the bitterling Rhodeus sinensis (Acheilognathidae) reveals genomic signatures associated with its mussel-dependent reproductive system.

Bitterlings (Acheilognathidae) exhibit a unique reproductive strategy characterized by symbiotic embryonic development inside the gill cavities of freshwater unionid mussels. Despite extensive ecological and physiological research on this system, genomic resources for bitterlings have remained limited, hindering comparative and evolutionary studies. Here, we present a high-quality, chromosome-level genome assembly for Rhodeus sinensis, a widely distributed bitterling species in the Korean Peninsula. By combining PacBio Continuous Long Read (CLR) sequencing, Illumina short reads, and Hi-C scaffolding, we generated a 0.77 Gb genome assembly with a scaffold N50 of 30.06 Mb. The final assembly comprises 24 chromosome-scale scaffolds, accounting for 98.3% of the assembled genome, with a BUSCO completeness score of 96.3% against the Actinopterygii_odb10. Comparative genomic analyses identified prominent expansions in gene families associated with alcohol metabolism, lipid catabolism, and oxidative stress responses. These genomic signatures of metabolic rewiring suggest a potential fuel flexibility, which may serve as a critical adaptive mechanism to mitigate the severe hypoxic stress encountered within the host mussel's gill environment. Ultimately, our chromosome-level genome assembly and findings provide a robust genomic foundation, contributing to a deeper understanding of the extreme physiological adaptations and unique life-history evolution within the Acheilognathidae.

Rhodeus sinensis

A chromosome-level genome assembly of Guimi No. 2 (Actinidia chinensis).

In this study, we report a high-quality chromosome-level genome assembly of Actinidia chinensis var. chinensis 'Guimi No. 2'. This cultivar, discovered in Guizhou karst ecosystems, exhibits resistance to Pseudomonas syringae pv. actinidiae (Psa). Using a combination of MGI short-read sequencing, PacBio HiFi long-read sequencing, and Hi-C technology, we generated a genome assembly of 608.43 Mb with a contig N50 of 20.70 Mb, and 99.70% of the assembly was successfully anchored onto 29 pseudochromosomes. The quality value (QV) and the LTR Assembly Index (LAI) of the assembled genome were 72.23 and 10.10. The BUSCO analysis indicated that the genome assembly and gene model prediction were 98.40% and 96.56% complete, respectively. A total of 251.15 Mb of repetitive sequences and 45,986 protein-coding genes were annotated. This genome assembly provides critical insights into A. chinensis's genomic architecture and serves as a foundational resource for elucidating disease resistance mechanisms against Psa, while enabling comparative phylogenomic studies across the Actinidia genus.

Actinidia

Draft genome assembly of the green-bronze dung beetle, Onthophagus orpheus.

Dung beetles (Coleoptera: Scarabaeinae) are ecologically important insects, yet genomic resources for this diverse lineage remain limited. Here, we present a high-quality genome assembly for Onthophagus orpheus, an understudied species that is abundant in urban forests in the eastern United States. The assembled genome is a scaffold-level assembly, with a high degree of genic completeness as assessed by Benchmarking Universal Single-Copy Ortholog (BUSCO) analyses, indicating robust representation of conserved protein-coding genes. Structural and functional annotation recovered a comprehensive gene set consistent with expectations for coleopteran genomes. This genome assembly provides an important resource for future work on the behavioral ecology and population genetics of Onthophagus orpheus, specifically, and Scarabaeidae more broadly.

Onthophagus

De Novo Whole Genome Assemblies of Unusual Case-Making Caddisflies (Trichoptera) Highlight Genomic Convergence in the Composition of the Major Silk Gene (h-fibroin).

Trichoptera (caddisflies) is one of the most species-rich orders of aquatic insects. Species of caddisflies cover a broad ecological diversity as exemplified by various uses of underwater silk secretions. Diversity of silk use generally aligns with the evolution of major caddisfly lineages, specifically at the subordinal level: Annulipalpia (retreat makers) and Integripalpia (cocoon and tube-case makers). However, silk use within suborders differs for a few exceptional species in these clades. In this study, we provide the first whole genome assemblies and annotations for two unusual Integripalpia species: Limnocentropus insolitus, whose hard tube-case is anchored to boulders by a rigid, elongated silken stalk, and Phryganopsyche brunnea which builds a "floppy" cylindrical case that lacks the typical robustness of tube-cases. Its texture rather resembles that of the flexible retreats built by Annulipalpia. Using the two high-quality genome assemblies, we identified and annotated the major silk gene, h-fibroin, and compared its amino acid composition across various groups, including retreat, cocoon, and tube-case makers. Our phylogenetic analysis confirmed the phylogenetic position of the two species in the tube-case-making clade. The major silk gene of L. insolitus shows a similar amino acid composition to other tube-case-making species. In contrast, the amino acid composition of P. brunnea resembles that of retreat-making species, in particular with regard to the high content of proline. This is consistent with the hypothesis that proline could be linked to enhanced extensibility of silk fibers. Taken together, our results underscore the role of silk genes in shaping the evolutionary ecology of retreat- and tube-case-making in caddisflies.

Animals

A high-quality chromosome-level genome assembly and annotation of the giant freshwater prawn (Macrobrachium rosenbergii).

The giant freshwater prawn, Macrobrachium rosenbergii, is native to Southeast Asia and is used in aquacultural practices worldwide. It is considered advantageous because of its rapid growth, high nutritional value, and economic benefits. As one of the three major freshwater aquaculture shrimp sources in China, a high-quality genome resource is of great significance for promoting the germplasm improvement of varieties. This study presents a high-quality chromosome-level genome assembly of M. rosenbergii that was generated by combining PacBio, MGI, and Hi-C reads. The assembled genome was 2.96 Gb in size, with a contig N50 of 0.64 Mb and a scaffold N50 of 55.76 Mb, which was positioned on 59 pseudo-chromosomes. The Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis for genome assembly reached 94.37%. In total, 27,111 protein-coding genes were identified, of which 25,470 were functionally annotated. These results provide a foundation for future research into adaptive evolution, genomics, and molecular breeding in M. rosenbergii.

Animals

Chromosome-level genome assembly of the horned turban snail Turbo cornutus.

The horned turban snail (Turbo cornutus) is an ecologically and economically important herbivorous gastropod inhabiting nearshore rocky reef habitats. T. cornutus represents a valuable coastal fishery resource in East Asia. Here, we present a chromosome-level genome assembly for T. cornutus generated using a combination of PacBio HiFi long-read and Illumina short-read sequencing and Hi-C scaffolding. The assembled genome spanned 1.93 Gb and was organized into 18 pseudo-chromosomes, representing 99.50% of the total assembly. The contig and scaffold N50 lengths were 41.02 Mb and 104.01 Mb, respectively, with repeat sequences constituting 59.07% of the genome. A total of 28,920 protein-coding genes were predicted, and genome completeness was assessed at 99.3% using the BUSCO mollusca_odb12 dataset. This chromosome-level genome assembly provides a reference for future studies on the biology of T. cornutus, the organization of the gastropod genome, and comparative genomics.

Animals

Colora: a Snakemake workflow for complete chromosome-scale de novo genome assembly.

MOTIVATION: De novo assembly creates reference genomes that underpin many modern biodiversity and conservation studies. Large numbers of new genomes are being assembled by labs around the world. To avoid duplication of efforts and variable data quality, we desire a best-practice assembly process, implemented as an automated portable workflow. RESULTS: Here, we present Colora, a Snakemake workflow that produces chromosome-scale de novo primary or phased genome assemblies complete with organelles using Pacific Biosciences HiFi, Hi-C, and optionally Oxford Nanopore Technologies reads as input. Colora is a user-friendly, versatile, and reproducible pipeline that is ready to use by researchers looking for an automated way to obtain high-quality de novo genome assemblies. AVAILABILITY AND IMPLEMENTATION: The source code of Colora is available on GitHub (https://github.com/LiaOb21/colora) and has been deposited in Zenodo under DOI https://doi.org/10.5281/zenodo.13321576. Colora is also available at the Snakemake Workflow Catalog (https://snakemake.github.io/snakemake-workflow-catalog/? usage=LiaOb21%2Fcolora).

Software

ONT-only genome assembly of a Korean male individual using a semen sample.

BACKGROUND: Long-read sequencing has enabled the generation of high-quality human genome assemblies, but many previous assemblies were based on blood-derived DNA and often relied on limited data types from a single sequencing strategy. OBJECTIVE: This study aimed to generate high-quality phased genome assemblies of a Korean individual using multiple independent long-read datasets produced from a single sequencing platform and to evaluate their utility for chromosome-scale assembly and variant detection. METHODS: Genomic DNA was extracted from a semen sample of a Korean male. Long-read, ultra-long-read, and chromatin conformation capture sequencing data were generated using Oxford Nanopore Technologies. These datasets were integrated to construct phased genome assemblies, followed by correction of noticeable phasing errors and assessment of assembly continuity, chromosomal representation, telomeric repeat recovery, and variant detection performance. RESULTS: The final phased assemblies spanned approximately 2.9 Gb and represented 23 pairs of chromosomes with an NG50 of 150 Mb. Telomeric repeats were detected at 36 and 37 of the 48 chromosomal ends in the two assemblies, indicating high end-to-end completeness. In addition, we successfully identified structural variants, including small variants. These results demonstrate that combining multiple Oxford Nanopore data types can produce highly continuous and informative phased human genome assemblies. CONCLUSIONS: We generated high-quality phased genome assemblies of a Korean individual using Oxford Nanopore long-read sequencing data derived from semen DNA. This publicly available genome resource will support broader applications of long-read sequencing in human genomics and variant analysis.

Humans

Chromosome-level genome assembly of Manglietia pachyphylla.

Manglietia pachyphylla, an endangered evergreen tree within the Magnoliaceae family, is renowned for its exceptional ornamental value in landscape horticulture. Despite its classification as a Category II nationally protected plant species in China, the genetic basis of its adaptive traits and conservation priorities remains poorly understood. To address this, we present the first chromosome-scale genome assembly of M. pachyphylla utilizing an integrated approach combining PacBio HiFi long-read and Hi-C chromosome conformation capture sequencing technologies. The assembled genome spans 2.15 Gb (contig N50 = 43.57 Mb), exhibiting a heterozygosity rate of 0.78% and repeat content of 78.64%, predominantly comprising long terminal repeat (LTR) retrotransposons (52.86%). Hi-C scaffolding anchored 99.57% of the assembly to 19 pseudochromosomes, achieving a BUSCO completeness score of 96.4%. Annotation revealed 42,505 putative protein-coding genes, with 84.46% of predicted genes were functionally annotated. Phylogenomic analysis positioned M. pachyphylla and Oyama sieboldii clustered together in a well-supported group. This high-contiguity genome assembly enables future investigations into adaptive evolution, functional genomics, and evidence-based conservation strategies for this endangered species.

Chromosomes, Plant

The genome sequence of a coronate scyphozoan jellyfish, Nausithoe racemosa (Komai, 1936) (Coronatae: Nausithoidae), and a metagenome-assembled genome of the associated cyanobacterium Moorena producens.

We present a genome assembly from a specimen of Nausithoe racemosa (coronate scyphozoan jellyfish; Cnidaria; Scyphozoa; Coronatae; Nausithoidae). The assembly contains two haplotypes with total lengths of 4 784.66 megabases and 4 868.20 megabases. Most of haplotype 1 (97.34%) is scaffolded into 20 chromosomal pseudomolecules. Haplotype 2 was assembled to scaffold level. The mitochondrial genome has also been assembled, with a length of 13.97 kilobases. From the metagenome data, we recovered one high-quality metagenome-assembled genome.

Coronatae

A chromosome-level, haplotype-resolved genome assembly for the barn owl, Tyto alba.

Recent advances in long-read sequencing have enabled near telomere-to-telomere (T2T) assemblies across diverse taxa. However, avian genomes remain challenging due to numerous microchromosomes, small, typically < 20Mb, DNA molecules that are gene-, GC-, and repeat-rich. As a consequence, microchromosomes are often missing from genome assemblies. Here, we present a chromosome-level, haplotype-resolved genome assembly for the Western barn owl (Tyto alba). Using a trio-binning strategy with Illumina parental reads combined with PacBio HiFi and Oxford Nanopore Technologies data, we generated two phased contig sets. These were scaffolded into 40 linkage groups using a linkage map. Comparative analyses identified unplaced HiFi scaffolds corresponding to microchromosomes, which we integrated into six additional microchromosomes using long reads information. The two assemblies present 46 chromosomes, matching the karyotype of the species. They exhibit strong synteny between parental haplotypes, except for a &#x223c;38 Mb complex region on chromosome 7 containing nested inversions. This high-quality reference provides a haplotype-resolved and chromosome-level genome for Strigiformes, enabling fine-scale studies of structural variation and avian genome evolution.

Tyto alba

Chromosome-level genome assembly of Ampulex clypecomplana Chen & Li (Hymenoptera: Ampulicidae).

Ampulex clypecomplana Chen & Li, 2010 (Hymenoptera: Ampulicidae) is an important predatory insect in Hymenoptera. However, molecular information about this predatory insect is currently limited. In this study, we employed ONT long-read sequencing, MGI-SEQ short-read sequencing, Hi-C sequencing and transcriptomic data to assemble the high-quality genome of A. clypecomplana. The genome assembly length was 338.43&#x2009;Mb, with a Scaffold N50 length of 19.05&#x2009;Mb. Our BUSCO analysis further confirmed the gene coverage completeness of the genome assembly to be 99.2%. Phylogenetic analysis indicated that A. clypecomplana appeared approximately 132 million years ago. We annotated 110.75&#x2009;Mb of repetitive sequences, accounting for 32.72% of the entire genome. In A. clypecomplana, we identified 180 gene expansions and 1029 genes that underwent contraction or loss. The high-quality genome of A. clypecomplana provides a valuable genetic resource for future research in evolution, molecular biology, and applied studies.

Animals

A telomere-to-telomere gap-free genome assembly of the endangered humphead wrasse (Cheilinus undulatus).

Humphead wrasse, Cheilinus undulatus, is an endangered fish species with high economic and ecological value as well as natural sex change from female to male, while sexual selection occurs in breeding aggregations. In our present study, we constructed the first gap-free telomere-to-telomere (T2T) genome assembly for humphead wrasse, by integration of PacBio HiFi, ONT Ultra-long and Hi-C sequencing techniques. With 99% of the entire sequences anchored into 24 chromosomes, this haplotypic genome assembly spans approximately 1.25&#x2009;Gb and presents a complete set of 48 telomeres and 24 centromeres. In terms of correctness (quality value QV: 53.447) and completeness (BUSCO score: 99.3%), this chromosome-scale assembly is indeed of high quality. We predicted 658.03&#x2009;Mb of repetitive sequences and annotated 26,609 protein-coding genes in the assembled genome. This high-quality T2T genome assembly not only facilitates the genetic conservation of humphead wrasse, but also offers fundamental genomic data for supporting in-depth investigations on functional genomics, genetic diversity, and selective breeding for this economically important teleost.

Animals

Revisiting the genome assembly of Lupinus species reveals differential diploidization after a shared whole-genome duplication.

Accurate genome assemblies are essential for comparative genomics, yet Hi-C-guided scaffolding can introduce structural errors that misrepresent chromosome architecture and bias evolutionary inferences. Here, we identified pervasive scaffolding errors-including artificial fusions, internal inversions, and incomplete contig mounting-in 2 previously published Lupinus genomes (L. cosentinii and L. digitatus) using a segmentation method based on long terminal repeat (LTR) retrotransposon density. We reassembled both genomes, producing chromosome-level references of 472.7 Mb (16 chromosomes) and 427.2 Mb (21 chromosomes), with BUSCO completeness >98.5%. Synteny validation and reapplication of LTR profiling confirmed that all prior errors were resolved. Using these corrected genomes together with 4 additional Lupinus species and 2 outgroup legumes, we investigated postpolyploid evolution. Synonymous substitution rate (Ks) analysis revealed a genus-specific whole-genome duplication (WGD) event (Ks = 0.17) shared by all 6 Lupinus species. The proportion of WGD-derived genes varied markedly, from 60% in L. digitatus to only 36% in L. mutabilis, indicating differential diploidization. While all species retained a core set of WGD duplicates enriched in cytoskeleton organization, ion transport, and defense responses, each exhibited lineage-specific functional trajectories: cell wall modification in L. cosentinii and L. digitatus, nitrogen metabolism in L. albus and L. angustifolius, flower development in L. luteus, and stress/lipid metabolism in L. mutabilis. Our corrected assemblies provide optimal references for Lupinus comparative genomics, and our findings demonstrate that a shared WGD event can lead to both conserved and highly divergent postpolyploid fates, likely underpinning adaptive diversification within the genus.

Lupinus

Chromosome-level genome assembly and annotation of Petunia hybrida.

Petunia hybrida is the world's most popular garden plant and is regarded as a supermodel for studying the biology associated with the Asterid clade, the largest of the two major groups of flowering plants. Unlike other Solanaceae, petunia has a base chromosome number of seven, not 12. This along with recombination suppression has previously hindered efforts to assemble its genome to chromosome level. Here we achieve a chromosome-level assembly for P. hybrida using a combination of short-read and long-read sequencing, optical mapping (Bionano) and Hi-C technologies. The resulting assembly spans 1253.6&#x2009;Mb with a BUSCO score of 99.8%. A total of 35,089 genes were predicted and of those 29,655 were functionally annotated. Syntenic regions between petunia, tomato and pepper were identified, highlighting rearrangements that have occurred since their divergence indicating that the 12 chromosomes of Solanaceae did not originate from whole genome duplication of an ancestral species with seven chromosomes like petunia. This assembly will enhance trait mapping efficiency and serve as a valuable resource for functional genomic studies.

Petunia

Chromosome-level genome assembly of Nothapodytes nimmoniana.

Nothapodytes nimmoniana is a plant species belonging to the genus Nothapodytes in the family Icacinaceae. This species holds significant medicinal value due to its camptothecin content. In this study, we present the first chromosome-level genome assembly of N. nimmoniana constructed using NGS, Hi-C, and HiFi sequencing technologies. The assembled genome spans 3.53&#x2009;Gb across 14 chromosomes, with an N50 length of 248.74&#x2009;Mb. Genome annotation revealed that repetitive sequences constitute 80.82% of the genome size, and 83,269 protein-coding genes were predicted. Additionally, 4,360,538&#x2009;bp of non-coding RNA were annotated. This genomic resource provides a foundation for further investigation into camptothecin biosynthesis pathways and plant phylogeny in N. nimmoniana.

Genome, Plant