Search PubMedSearch

SEARCH · Search PubMed

Results for “Hi-C scaffolding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Chromosomal-level genome assembly of Trypanosoma carassii, the etiologic agent of a recent outbreak of trypanosomiasis in cage-cultured large yellow croaker (Larimichthys crocea) in China.

Trypanosoma carassii, a typical freshwater fish trypanosome, has recently been identified as the etiological agent of a trypanosomiasis outbreak in cage-cultured large yellow croaker (Larimichthys crocea) in China and has been designated as T. c. larimichthys. To date, publicly available genomic data for trypanosomes have been limited to terrestrial species, particularly those of medical importance. Here, we present a chromosome-level genome assembly of T. carassii, the first genome of an aquatic trypanosome, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. A preliminary genome survey based on Illumina sequencing data estimated the genome size at 56.38 Mb with a heterozygosity of 1.17%. The final assembled genome spans 48.55 Mb, with contig N50 and scaffold N50 values of 139.15 Kb, and achieves 100.00% BUSCO completeness. Hi-C data resolved the assembly into 34 chromosomes and 9 unanchored scaffolds. Repetitive elements account for 53.29% of the genome (approximately 25.87 Mb). A total of 11,584 protein-coding genes were predicted, 95.36% of which were functionally annotated. Synonymous substitution rates analysis of paralogous genes indicates a recent burst of gene duplication, which likely corresponds to a whole-genome duplications. This high-quality genome assembly provides invaluable resources for understanding the evolution and host adaptation of aquatic trypanosomes.

Animals

A chromosome-level genome assembly of Coffea arabica L. var. 'Kona Typica'.

Coffea arabica L. var. 'Kona Typica' is renowned for its premium cup quality, but its vulnerability to pests and diseases limits production. To accelerate cultivar improvement, we generated a chromosome-level genome assembly of 'Kona Typica' using PacBio HiFi sequencing and Hi-C scaffolding technology. The final assembly spans 1.13 Gb, with a scaffold N50 of 50.50 Mb, organized into 22 chromosomes. BUSCO assessment indicated a high completeness at 99.1%. We annotated 65,458 protein-coding genes and identified 1,073,545 interspersed repeats, accounting for 65.16% of the genome. Analysis of transposon insertion ages revealed that most long terminal repeat retrotransposons proliferated after the polyploidization event. This high-quality genome assembly of 'Kona Typica' provides a valuable resource for exploring coffee genomic evolution and genetic mechanisms of complex traits, facilitating genomics studies and the development of improved coffee cultivars with enhanced disease resistance and quality traits.

Coffea

Chromosome-level genome assembly of Manglietia pachyphylla.

Manglietia pachyphylla, an endangered evergreen tree within the Magnoliaceae family, is renowned for its exceptional ornamental value in landscape horticulture. Despite its classification as a Category II nationally protected plant species in China, the genetic basis of its adaptive traits and conservation priorities remains poorly understood. To address this, we present the first chromosome-scale genome assembly of M. pachyphylla utilizing an integrated approach combining PacBio HiFi long-read and Hi-C chromosome conformation capture sequencing technologies. The assembled genome spans 2.15 Gb (contig N50 = 43.57 Mb), exhibiting a heterozygosity rate of 0.78% and repeat content of 78.64%, predominantly comprising long terminal repeat (LTR) retrotransposons (52.86%). Hi-C scaffolding anchored 99.57% of the assembly to 19 pseudochromosomes, achieving a BUSCO completeness score of 96.4%. Annotation revealed 42,505 putative protein-coding genes, with 84.46% of predicted genes were functionally annotated. Phylogenomic analysis positioned M. pachyphylla and Oyama sieboldii clustered together in a well-supported group. This high-contiguity genome assembly enables future investigations into adaptive evolution, functional genomics, and evidence-based conservation strategies for this endangered species.

Chromosomes, Plant

High-Resolution Chromosome-Level Genome Assembly and Annotation of Triplophysa stewarti, an Endemic Plateau Loach from the Qinghai-Tibet Plateau.

The bottom-dwelling fish Triplophysa stewarti, endemic to the Qinghai-Tibet Plateau, is a valuable model for studying high-altitude adaptation in aquatic ecosystems. However, the lack of a high-quality reference genome has hindered comparative genomic and evolutionary studies within this genus. Here, we present a chromosome-level genome assembly for T. stewarti, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding. The 697.9 Mb assembly is highly continuous (scaffold N50 of 253.58 Mb) and encompasses 25 chromosomes, representing 92.65% of the genome. BUSCO analysis indicated a 98.4% completeness, supporting the high quality of the assembly. We annotated 28,009 protein-coding genes, with 97.04% being functionally assigned across multiple databases (NR, UniProt, KEGG, GO, Pfam and InterPro). Additionally, repetitive elements constituted 42.47% of the genome, and we identified 52,709 non-coding RNAs. This high-quality reference genome provides a fundamental resource for exploring the adaptive evolution, population structure, and conservation genetics of T. stewarti and related species on the Qinghai-Tibet Plateau.

Animals

Chromosome-level genome assembly of the small-sized Taihang donkey (Equus asinus).

China harbors a rich diversity of donkey breeds, with small-sized donkeys (<110&#x2009;cm) representing a largely underexplored group. Here, we present the first high-quality, chromosome-level genome assembly of a small-sized donkey, generated using PacBio HiFi sequencing (286.7&#x2009;Gb), Hi-C scaffolding (240.47&#x2009;Gb), and annotated with RNA-seq data. The final assembly has a total length of 2.7&#x2009;Gb and comprises 32 chromosomes (including both X and Y chromosomes), in which five chromosomes were fully assembled without gaps. It possesses a scaffold N50 of 106.70&#x2009;Mb and 84 contigs (contig N50&#x2009;=&#x2009;63.60&#x2009;Mb), and captures 99.2% of BUSCO genes. The assembly achieved a consensus quality value (QV) of 77.44, corresponding to an extremely low base-level error rate, indicating exceptional nucleotide accuracy. This high-quality genome provides a valuable resource for investigating genetic variation, adaptive evolution, and domestication processes in small-sized donkeys, and will facilitate the conservation and sustainable utilization of rich donkey genetic resources in China.

Animals

Revisiting the genome assembly of Lupinus species reveals differential diploidization after a shared whole-genome duplication.

Accurate genome assemblies are essential for comparative genomics, yet Hi-C-guided scaffolding can introduce structural errors that misrepresent chromosome architecture and bias evolutionary inferences. Here, we identified pervasive scaffolding errors-including artificial fusions, internal inversions, and incomplete contig mounting-in 2 previously published Lupinus genomes (L. cosentinii and L. digitatus) using a segmentation method based on long terminal repeat (LTR) retrotransposon density. We reassembled both genomes, producing chromosome-level references of 472.7 Mb (16 chromosomes) and 427.2 Mb (21 chromosomes), with BUSCO completeness >98.5%. Synteny validation and reapplication of LTR profiling confirmed that all prior errors were resolved. Using these corrected genomes together with 4 additional Lupinus species and 2 outgroup legumes, we investigated postpolyploid evolution. Synonymous substitution rate (Ks) analysis revealed a genus-specific whole-genome duplication (WGD) event (Ks = 0.17) shared by all 6 Lupinus species. The proportion of WGD-derived genes varied markedly, from 60% in L. digitatus to only 36% in L. mutabilis, indicating differential diploidization. While all species retained a core set of WGD duplicates enriched in cytoskeleton organization, ion transport, and defense responses, each exhibited lineage-specific functional trajectories: cell wall modification in L. cosentinii and L. digitatus, nitrogen metabolism in L. albus and L. angustifolius, flower development in L. luteus, and stress/lipid metabolism in L. mutabilis. Our corrected assemblies provide optimal references for Lupinus comparative genomics, and our findings demonstrate that a shared WGD event can lead to both conserved and highly divergent postpolyploid fates, likely underpinning adaptive diversification within the genus.

Lupinus

Chromosome-level assembly and annotation of the Jaguar (Panthera onca) genome.

OBJECTIVES: The Jaguar (Panthera onca) is a large cat species native to the Americas. Despite being successful predators, jaguar populations have declined due to habitat loss. Genome resources can help in conservation efforts as well as in understanding the interesting biology of these Felids. Beside contiguity, a well annotated reference genome provides contextual information for variants that will benefit the design of appropriate conservation programs. DATA DESCRIPTION: We sequenced material from two individuals using a combination of ONT reads and Illumina PE. The resulting nuclear genome assembly has a larger contig N50 (48.04&#xa0;Mb) compared with the existing annotated chromosome-level assembly published by the DNA Zoo project. Using public Hi-C data, we obtained an improved chromosome-level assembly of the Jaguar genome (mPanOnc3.5) with larger contigs, 99.85% of the sequence assigned to chromosomes and 25,267 protein coding genes annotated. Overall, this improved assembly provides a better reference to study this threatened species.

Animals

Characterization of a draft chromosome-scale genome assembly for the mutton snapper, Lutjanus analis.

BACKGROUND: The mutton snapper (Lutjanus analis) is a reef fish commonly found in tropical waters of the Western Atlantic Ocean. Genomic studies of this species are needed to support conservation efforts and breeding programs. OBJECTIVE: Here, we report the development of a chromosome-scale reference assembly for the mutton snapper and conduct an initial comparative genomic analysis with other lutjanids. METHODS: The genome of one mutton snapper specimen was sequenced using PAC-Bio HiFi long reads and Illumina short reads. Contigs and scaffolds were assembled in the Flye pipeline and anchored using Hi-C proximity guided assembly. Gene prediction and functional annotations were obtained in AUGUSTUS and eggNOG-mapper, respectively. The mutton snapper genome was compared to those of other lutjanids to infer gene family evolution and chromosome synteny conservation. RESULTS: Assembly and polishing yielded 946 contigs and 926 scaffolds (N50 of 3.16&#xa0;Mb, complete BUSCO score 98.1%) that were anchored using Hi-C scaffolding in 24 draft chromosomes. The anchored assembly featured a N50 of 42.47&#xa0;Mb and contained 97.6% of the unanchored assembly length. The 24 mutton snapper chromosomes showed a one-to-one syntenic relationship with their counterparts in medaka, and other Lutjanids. AUGUSTUS predicted 29,023 genes, 24,335 of which (83.85%) could be functionally annotated. Gene family evolution analysis revealed 1,014 significantly expanded or contracted hierarchical ortholog groups in mutton snapper. Expansions and contractions were linked to several biological functions including growth, oocyte maturation, and response to exogenous stressors. CONCLUSION: The draft genome will be a valuable tool for forthcoming applied genomic studies of mutton snapper.

Animals

Chromosome-level genome assembly of Sinocyclocheilus jii based on PacBio HiFi and Hi-C sequencing.

Sinocyclocheilus jii, a cavefish species endemic to China, belongs to the genus Sinocyclocheilus within the family Cyprinidae. Species within this genus exhibit significant morphological differentiation, making it not only the most species-rich genus within Cyprinidae in China but also the most diverse group of cavefishes worldwide. However, the limited availability of genomic resources has limited investigations into the genetic basis of trait variations, phylogenetic relationships, and adaptive evolution in this genus. In this study, we assembled a chromosome-level reference genome for S. jii by integrating PacBio HiFi long reads, Illumina short reads, and Hi-C sequencing data. Flow cytometry was used to estimate the genome size prior to assembly, providing a key step in technical validation. The final genome assembly spans 1.75&#x2009;Gb with a contig N50 of 35.0&#x2009;Mb. Using Hi-C sequencing data, the assembled scaffolds were successfully anchored to 50 chromosomes. The completeness of the chromosome-level assembly was estimated at 98.9% by BUSCO analysis. Genome annotation identified 855.5&#x2009;Mb of repetitive sequences and predicted a total of 52,867 protein-coding genes, of which 51,932 genes were functionally annotated. This study presents a high-quality chromosome-level genome assembly and annotation of S. jii, providing a fundamental genomic resource for future phylogenetic and evolutionary studies.

Animals

Chromosome-scale assembly with improved annotation provides insights into breed-wide genomic structure and diversity in domestic cats.

INTRODUCTION: Comprehensive genomic resources offer insights into biological features, including traits/disease-related genetic loci. The current reference genome assembly for the domestic cat (Felis catus), Felis_Catus_9.0 (felCat9), derived from sequences of the Abyssinian cat, may inadequately represent the general cat population, limiting the extent of deducible genetic variations. OBJECTIVES: The goal was to develop Anicom American Shorthair 1.0 (AnAms1.0), a reference-grade chromosome-scale cat genome assembly. METHODS: In contrast to prior assemblies relying on Abyssinian cat sequences, AnAms1.0 was constructed from the sequences of more popular American Shorthair breed, which is related to more breeds than the Abyssinian cat. By combining advanced genomics technologies, including PacBio long-read sequencing and Hi-C- and optical mapping data-based sequence scaffolding, we compared AnAms1.0 to existing Felidae genome assemblies (20 scaffolds, scaffolds N50&#xa0;>&#xa0;150 Mbp). Homology-based and ab initio gene annotation through Iso-Seq and RNA-Seq was used to identify new coding genes and splice variants. RESULTS: AnAms1.0 demonstrated superior contiguity and accuracy than existing Felidae genome assemblies. Using AnAms1.0, we identified over 1.5 thousand structural variants and 29 million repetitions compared to felCat9. Additionally, we identified > 1,600 novel protein-coding genes. Notably, olfactory receptor structural variants and cardiomyopathy-related variants were identified. CONCLUSION: AnAms1.0 facilitates the discovery of novel genes related to normal and disease phenotypes in domestic cats. The analyzed data are publicly accessible on Cats-I (https://cat.annotation.jp/), which we established as a platform for accumulating and sharing genomic resources to discover novel genetic traits and advance veterinary medicine.

Animals

A gap-free, telomere-to-telomere chromosome-scale genome assembly of the mangrove red snapper, Lutjanus argentimaculatus.

The mangrove red snapper (Lutjanus argentimaculatus) is a commercially important marine fish species in the Indo-Pacific region. Despite its significant economic value for aquaculture, existing genomic resources remain fragmented, limiting the advancement of molecular breeding and functional genomic studies. Here, we present a gap-free, telomere-to-telomere (T2T) genome assembly of L. argentimaculatus, generated using a hybrid approach combining PacBio HiFi, Oxford Nanopore ultra-long reads and Hi-C technology. The resulting assembly comprises exactly 24 scaffolds spanning 1.03&#x2009;Gb, perfectly matching the haploid chromosome number with a contig N50 of 46.17&#x2009;Mb. Notably, this assembly resolves all physical gaps present in previous versions, achieving a BUSCO completeness score of 98.2%. Comprehensive genome annotation successfully predicted 23,167 protein-coding genes. Among these, 22,067 genes (95.25%) were functionally annotated across major public databases, including eggNOG, InterPro, and Swiss-Prot. Furthermore, structural analysis successfully identified 19 telomeres and 20 centromeres, validating the chromosomal integrity. This high-fidelity, gap-free reference genome provides a robust foundation for comparative genomics, population genetics, and the genetic improvement of Lutjanidae species.

Animals

Assembling genomes of non-model plants: A case study with evolutionary insights from Ranunculus (Ranunculaceae).

Whereas genome sequencing and assembly technologies are improving, cost can still be prohibitive for plant species with large, complex genomes. As a consequence, genomics work on some taxa in evolutionarily pivotal positions in the vascular plant tree of life has been hampered. The species-rich genus Ranunculus (Ranunculaceae) is an important angiosperm group for the study of polyploidy, apomixis, and reticulate evolution. However, neither mitochondrial nor high-quality nuclear genome sequences are available. This limits phylogenomic, functional, and taxonomic analyses thus far. Here, we tested Illumina short-read, Oxford Nanopore Technology (ONT) and PacBio (HiFi) long-read, and hybrid-read assembly strategies. We sequenced the diploid progenitor species R. cassubicifolius (R.&#x2009;auricomus species complex) and selected the best assemblies in terms of completeness, contiguity, and quality scores. We first assembled the plastome (156&#x2009;kbp, 85 genes) and mitogenome (1.18&#x2009;Mbp, 40 genes) sequences using Illumina and Illumina-PacBio-hybrid strategies, respectively. We also present an updated plastome and the first mitogenome phylogeny of Ranunculaceae, including studies of gene loss (e.g., infA, ycf15, or rps) with evolutionary implications. For the nuclear genome sequence, we favored a PacBio-based assembly polished three times with filtered short reads and subsequently scaffolded into eight pseudochromosomes by chromatin conformation data (Hi-C). We obtained a haploid genome sequence of 2.69&#x2009;Gbp, with 94.1% complete BUSCO genes found and 35&#x2009;482 annotated genes, and inferred ancient gene duplications compared to existing Ranunculales genomes. The genomic information presented here will enable advanced evolutionary-functional analyses for the species complex, but also for the genus and beyond Ranunculaceae.

Ranunculus

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55&#x2009;Gb, with a scaffold N50 of 93.38&#x2009;Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant

Chromosome-level genome assembly of the Vermilion Snapper (Rhomboplites aurorubens).

Vermilion Snapper (Rhomboplites aurorubens, Lutjanidae) inhabits deep waters (20-300&#x2009;m) from North America to Brazil and supports significant commercial and recreational fisheries. Despite its economic importance, the understanding of its basic biology remains limited. Classified as Vulnerable on the Red List due to overfishing, populations have declined by over 30% in recent generations. We assembled and annotated the first chromosome-scale genome of this species by combining PacBio long reads, Illumina short reads, and Hi-C data. The resulting assembly is 987.5 Mbp, with a scaffold N50 size of 41.3 Mbp, and includes 135 contigs clustered and ordered onto 24 chromosomes with 34,496 predicted genes. The high-quality assembly and annotation contained about 98% complete and single-copy BUSCO genes. It is the most complete, chromosome-level genome assembly of an Atlantic snapper to date. The genome assembly and supporting data are valuable tools for ecological and comparative genomics studies of snappers and other valuable commercial species within the family.

Chromosomes

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62&#x2009;Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52&#x2009;Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals

Chromosome level genome assembly and full-length transcriptome of blacktip trevally (Caranx heberi).

Caranx heberi (Bennett, 1830) commonly known as the blacktip trevally belongs to the family Carangidae and is a potential brackishwater aquaculture species. However, the limited genomic resources are hindering the efforts to study its genetic traits and their molecular basis. To bridge this gap, we generated a high-quality reference genome employing multiple sequencing strategies including PacBio Hifi reads (135x), Illumina short reads (150x), and Hi-C chromosome conformation capturing (180x). The high-quality genome assembly consisted of 159 scaffolds summing to 618.71&#x2009;Mb and an N50 value of 26.72&#x2009;Mb. Among these, 24 chromosome level scaffolds covered 97.5% of the total assembly. The genome contained 20.94% of repeat elements and 30,354 protein encoding genes. In addition, full-length transcriptomes were generated using the PacBio IsoSeq approach from seven tissues (gill, kidney, liver, muscle, heart, spleen, and intestine). The comprehensive genomic and transcriptomic resources developed in this study will facilitate the domestication and aquaculture development of C. heberi, as well as support research on its nutritional potential, ecological adaptations, and evolutionary biology.

Animals

Chromosome-level genome assembly of Ampulex clypecomplana Chen & Li (Hymenoptera: Ampulicidae).

Ampulex clypecomplana Chen & Li, 2010 (Hymenoptera: Ampulicidae) is an important predatory insect in Hymenoptera. However, molecular information about this predatory insect is currently limited. In this study, we employed ONT long-read sequencing, MGI-SEQ short-read sequencing, Hi-C sequencing and transcriptomic data to assemble the high-quality genome of A. clypecomplana. The genome assembly length was 338.43&#x2009;Mb, with a Scaffold N50 length of 19.05&#x2009;Mb. Our BUSCO analysis further confirmed the gene coverage completeness of the genome assembly to be 99.2%. Phylogenetic analysis indicated that A. clypecomplana appeared approximately 132 million years ago. We annotated 110.75&#x2009;Mb of repetitive sequences, accounting for 32.72% of the entire genome. In A. clypecomplana, we identified 180 gene expansions and 1029 genes that underwent contraction or loss. The high-quality genome of A. clypecomplana provides a valuable genetic resource for future research in evolution, molecular biology, and applied studies.

Animals

A chromosome-level genome of the Nicobar pigeon, Caloenas nicobarica.

The Nicobar pigeon (Caloenas nicobarica), the closest living relative of the extinct Dodo (Raphus cucullatus), is endemic to Southeast Asia with a fragmented distribution across numerous small islands. It suffers from habitat loss, hunting, and predation from invasive species, resulting in its classification as Near Threatened by the International Union for the Conservation of Nature. We have generated a haplotype-resolved and chromosome-level genome assembly of the Nicobar pigeon using a combination of PacBio HiFi long-read sequencing and Arima Hi-C chromatin interaction mapping. This assembly includes two haplotypes, each spanning approximately 1.2 Gb. Haplotype 1 has a contig N50 of 25.2&#xa0;Mb and a scaffold N50 of 79.7&#xa0;Mb, whereas haplotype 2 has a contig N50 of 24.7&#xa0;Mb and a scaffold N50 of 107.9&#xa0;Mb. As the first high-quality genome assembly of any bird in the Columbidae Indo-Pacific clade, this resource provides valuable insights for phylogenetic studies. Furthermore, the phylogenetic proximity of the Nicobar pigeon to the Dodo (R. cucullatus) and the Rodrigues Solitaire (Pezophaps solitaria) offers a unique opportunity to study these extinct species, making this assembly a critical resource for evolutionary studies. It also offers a unique model for studying genetic diversity, adaptation, and speciation in island environments. This genomic resource will not only enhance our understanding of the evolutionary history of the Nicobar pigeon but also serve as a valuable tool for future conservation efforts aimed at preserving this unique species and its fragile island ecosystem.

Animals