Search PubMedSearch

SEARCH · Search PubMed

Results for “Genome assembly”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Microbiome Datahub: an open-access platform integrating environmental metadata, taxonomy, and functional annotation for comprehensive metagenome-assembled genome datasets.

BACKGROUND: Metagenome-assembled genomes (MAGs) provide crucial insights into the genomic diversity of uncultured microbes. However, MAG datasets deposited in public repositories such as INSDC are often difficult to reuse due to heterogeneous quality, inconsistent taxonomic and functional annotations, and insufficiently curated environmental metadata. While secondary MAG databases such as MGnify, IMG/M, and SPIRE provide standardized resources, they reconstruct MAGs de novo from public metagenomic reads and therefore do not represent the original MAGs reported in publications. RESULTS: To address this gap, we developed Microbiome Datahub, an open-access platform that systematically aggregates and re-annotates original MAGs from INSDC. We collected 214,427 MAGs, predicted genes by DFAST, performed quality assessment with CheckM, standardized taxonomic assignments with GTDB-Tk, inferred 27 phenotypic traits using Bac2Feature, assigned proteins to MBGD ortholog clusters and KEGG Orthology IDs using PZLAST, and annotated environmental metadata with the Metagenome and Microbes Environmental Ontology. Across these MAGs, the average completeness was 80.5% and contamination 1.8%; notably, the most frequent values were&#x2009;>95% completeness and&#x2009;<1% contamination, indicating that the majority of MAGs are of high quality. Comparative analyses showed that Microbiome Datahub provides phylogenetically and environmentally diverse MAGs: while the majority originated from vertebrate gut environments, a substantial number were also recovered from other habitats such as groundwater, including nearly 10,000 MAGs from the Patescibacteria. Inference of 27 phenotypic traits, including optimum growth temperature, further revealed ecological differentiation across phyla. Protein clustering revealed 56 million identity 40% clusters, with the majority unique compared with MGnify and GlobDB, and&#x2009;~19% of proteins unassigned to MBGD ortholog clusters, underscoring their novelty. CONCLUSIONS: Microbiome Datahub integrates MAG genome sequences, gene and protein predictions, quality metrics, environmental and taxonomic annotations, ortholog cluster assignments, and phenotype predictions, all accessible via a web interface, API, and bulk downloads. By combining original MAGs with curated metadata and functional annotations, Microbiome Datahub constitutes a comprehensive and reusable resource that will accelerate microbiome and microbial genomics research. Video Abstract.

Metagenome

Chromosome-level genome assembly and annotation of Spinibarbus caldwelli.

Spinibarbus caldwelli is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in southern China, the lack of genomic resources has hampered the genetic breeding and conservation. Here, we release a chromosome-level genome assembly for S. caldwelli using PacBio HiFi long-reads, Illumina short-reads, and Hi-C sequencing data. The final genome assembly is 1.77&#x2009;Gb in size, with a contig N50 of 24.27&#x2009;Mb. Using Hi-C scaffolding, 99.14% of the contigs were successfully anchored to 50 chromosomes, resulting in a scaffold N50 of 35.29&#x2009;Mb. The final genome assembly shows a BUSCO completeness of 98.27%. The assembled genome contains 49.41% repetitive sequences and 51,505 predicted genes, 90.83% of which have been functionally annotated. This genome provides a genetic basis for S. caldwelli, facilitating the exploration of Cyprinid phylogeny, genetic improvement, and conservation efforts.

Animals

Chromosome-level genome assembly of a cosmopolitan marine harmful algal bloom diatom species Chaetoceros socialis (Chaetocerotaceae).

Chaetoceros socialis is a cosmopolitan diatom species that is crucial for maintaining marine ecosystem structure and driving elemental cycles. C. socialis can form harmful algal blooms (HABs) that may cause a negative impact on the marine ecosystems. Whole-genome information for C. socialis is still unavailable, which may hinder more targeted studies on its ecological adaptive responses and evolutionary drivers. To address this gap, we employed cutting-edge genomic technologies including PacBio single-molecule real-time (SMRT) sequencing and high-throughput chromatin conformation capture (Hi-C) to achieve the first chromosome-level genome assembly of C. socialis. The assembled genome is 60.22&#x2009;Mb in size with a scaffold N50 of 7.81&#x2009;Mb and has been anchored to eight pseudochromosomes. A total of 13,378 protein-coding genes were predicted, of which 12,069 (90.22%) were functionally annotated. This high-quality genomic resource provides a fundamental data platform for systematically elucidating the ecological adaptation mechanisms of C. socialis.

Chromosomes

Agptools: a utility suite for editing genome assemblies.

SUMMARY: The AGP format is a tab-separated table format describing how components of a genome assembly fit together. A standard submission format for genome assemblies is a fasta file giving the sequence of contigs along with an AGP file showing how these components are assembled into larger pieces like scaffolds or chromosomes. For this reason, many scaffolding software pipelines output assemblies in this format. However, although many programs for assembling and scaffolding genomes read and write this format, there is currently no published software for making edits to AGP files when performing assembly curation. We present agptools, a suite of command-line programs that can perform common operations on AGP files, such as breaking and joining sequences, inverting pieces of assembly components, assembling contigs into larger sequences based on an AGP file, and transforming between coordinate systems of different assembly layouts. Additionally, agptools includes an API that writers of other software packages can use to read, write, and manipulate AGP files within their own programs. AVAILABILITY AND IMPLEMENTATION: Source code and binaries freely available for download at https://github.com/WarrenLab/agptools, implemented in Python and supported on all operating systems.

Software

Chromosome-level genome assembly of Cheilinus chlorourus (Bloch, 1791) (Perciformes: Labridae).

In the classification of marine fish, the Labridae family ranks second in terms of species diversity and plays a vital role in coral reef ecosystems, comprising over 600 species across 82 genera. Despite its significance for ecological and evolutionary studies, genomic research on this group has lagged, resulting in a shortage of data, particularly regarding high-quality chromosome-level genome assemblies. To address this gap, this study focused on Cheilinus chlorourus from the Labridae family and successfully achieved a chromosome-level genome assembly. By integrating Illumina, PacBio, and Hi-C sequencing data, we assembled a genome measuring 940.36&#x2009;Mb, with 926.86&#x2009;Mb (98.56%) of the gene assembly organized into 21 chromosomes. A total of 29,213 protein-coding genes (PCGs) were identified, and 79.93% of these genes were functionally annotated. With this high-quality genome assembly, future investigations into the functional genomics and ecology of C. chlorourus will have a solid scientific foundation.

Animals

Chromosome-level genome assembly of Qihe gibel carp.

Qihe gibel carp (Carassius gibelio var. Qihe) is a local population of natural gynogenetic amphitriploid (AAABBB) Carassius gibelio, and has high nutritional and economic value. In this study, we assemble a high-quality chromosome-level genome of Qihe gibel carp through DNBSEQ, PacBio HiFi, and Hi-C sequencing data. The resulting assembly consisted of 350 contigs with the full length of 1.607&#x2009;Gb and 96.21% (1.515&#x2009;Gb) of the assembled genome was successfully anchored to 50 chromosomes, with a contig N50 of 28.97&#x2009;Mb and a scaffold N50 of 29.84&#x2009;Mb. Repeated sequences accounting for 43.72% (732.494&#x2009;Mb) of the total were also identified, and gene prediction revealed 46,131 protein-coding genes with an annotation ratio of 96.48%. Furthermore, Benchmarking Universal Single-Copy Orthologue (BUSCO) analysis demonstrated that the genome assembly achieved high completeness, with a score of 97.66%. This high-quality chromosome-level genome lays the foundation for molecular biology research as well as molecular breeding and evolutionary studies of Qihe gibel carp in the future.

Animals

A high-quality draft genome assembly of Johnsongrass illuminates relationships between polyploidization, crop-wild hybridization, and reproductive biology.

Johnsongrass [Sorghum halepense (L.) Pers.] is an allopolyploid, rhizomatous, perennial grass species and one of the most troublesome weeds in global agriculture. We assembled the first Johnsongrass genome to clarify poorly understood genetic factors influencing variable rates of crop-wild hybridization with cultivated sorghum [S. bicolor (L.) Moench]. The draft genome assembly has a total size of 3.26 Gb and BUSCO completeness of 95.3%. We also report the first evolutionary analysis of INHIBITION OF ALIEN POLLEN (IAP), the only known cross-(in)compatibility locus in the genus. Our results reveal an evolutionary history of genome instability, including the loss of distinct parental subgenomes, and suggest that Nebraska accession 'J-37,' the genome donor, is a segmental allotetraploid that may function as a diploid or aneuploid during meiosis. Genome instability could explain observations of variable ploidies in Johnsongrass and facilitate ongoing hybridization with sorghum where gamete ploidies and IAP alleles match. Given this information, we provide a suggested research framework for studying evolution and gene expression in the Sorghum genus where crop-wild hybridization occurs and for predicting the potential for hybridization between specific crossing partners. Collectively, this work will bolster efforts to study and manage reproductive biology in other crop-wild polyploid complexes.

Sorghum

Chromosome-level genome assembly of the longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae).

The longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae) is a widely distributed wood-boring pest of conifers. Here, we assembled a chromosome-level genome of A. rusticus using Illumina, Oxford Nanopore, and Hi-C sequencing technologies. The assembled genome is 1180.40&#x2009;Mb, with a scaffold N50 of 125.01&#x2009;Mb, and BUSCO completeness of 93.6%. All contigs were assembled into ten pseudo-chromosomes. The genome contains 69.87% repeat sequences. We identify 18, 377 protein-coding genes in the genome, of which 11,368 were functionally annotated. This genome provides a valuable resource for understanding the ecology, genetics, and evolution of A. rusticus, as well as for controlling wood-boring pests.

Animals

Chromosome-Level Genome Assembly of Solanum carolinense.

Horsenettle (Solanum carolinense L.) is a noxious weed widely distributed across North America and increasingly invasive in other regions. Its strong environmental adaptability, complex defense strategies, and distinctive reproductive traits make it an important model for studying plant-herbivore coevolution. However, the absence of high-quality genomic resources has limited deeper investigation into its adaptive evolutionary mechanisms. In this study, we generated a chromosome-level reference genome assembly for S. carolinense using an integrated approach combining PacBio HiFi long-read sequencing, Illumina second-generation sequencing, and Hi-C chromatin interaction scaffolding. The final genome assembly had a total length of 915.40 Mb, with a contig N50 of 51.06 Mb and a scaffold N50 of 73.17 Mb; 96.05% of the sequences were successfully anchored onto 12 pseudochromosomes. The genome was characterized by a high proportion of repetitive sequences (73.64%) and substantial heterozygosity (1.13%), consistent with a highly repetitive and moderately high heterozygous genome. BUSCO analysis indicated that the chromosome-level genome assembly of S. carolinense reached a completeness score of 94.8%. A total of 32,206 protein-coding genes were annotated, of which 97.95% received functional annotations. The evaluation of the annotated protein-coding gene set returned a completeness value of 94.9%. This reference genome provides a valuable resource for advancing research on the adaptive evolution of weedy Solanaceae species, supports the development of more effective management strategies for this troublesome species, and offers a technical reference for assembling other highly heterozygous weed genomes.

Solanum carolinense

Chromosomal-level genome assembly of Trypanosoma carassii, the etiologic agent of a recent outbreak of trypanosomiasis in cage-cultured large yellow croaker (Larimichthys crocea) in China.

Trypanosoma carassii, a typical freshwater fish trypanosome, has recently been identified as the etiological agent of a trypanosomiasis outbreak in cage-cultured large yellow croaker (Larimichthys crocea) in China and has been designated as T. c. larimichthys. To date, publicly available genomic data for trypanosomes have been limited to terrestrial species, particularly those of medical importance. Here, we present a chromosome-level genome assembly of T. carassii, the first genome of an aquatic trypanosome, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. A preliminary genome survey based on Illumina sequencing data estimated the genome size at 56.38&#x2009;Mb with a heterozygosity of 1.17%. The final assembled genome spans 48.55&#x2009;Mb, with contig N50 and scaffold N50 values of 139.15 Kb, and achieves 100.00% BUSCO completeness. Hi-C data resolved the assembly into 34 chromosomes and 9 unanchored scaffolds. Repetitive elements account for 53.29% of the genome (approximately 25.87&#x2009;Mb). A total of 11,584 protein-coding genes were predicted, 95.36% of which were functionally annotated. Synonymous substitution rates analysis of paralogous genes indicates a recent burst of gene duplication, which likely corresponds to a whole-genome duplications. This high-quality genome assembly provides invaluable resources for understanding the evolution and host adaptation of aquatic trypanosomes.

Animals

First chromosome-level genome assembly of the colonial chordate model Botryllus schlosseri (Tunicata).

BACKGROUND: Botryllus schlosseri (Tunicata) is a colonial, laboratory model tunicate recognized for its remarkable developmental diversity, its regenerative abilities, and its peculiar genetically determined allorecognition system governed by a polymorphic locus controlling chimerism and cell parasitism. RESULTS: We report the first chromosome-level genome assembly of B. schlosseri subclade A1. By integrating long and short reads with Hi-C scaffolding, we produced both a phased diploid genome assembly and a conventional collapsed consensus sequence of 533 Mb. Of this total length, 96% belonged to 16 chromosome-scale scaffolds, with a BUSCO completeness score of 91.4%. We then compared our assembly with other high-quality tunicate genomes, revealing some synteny conservation but also extensive genomic rearrangements and a general loss of colinearity. CONCLUSIONS: The chromosome-level resolution of this assembly enhances our understanding of genome organization in colonial modular organisms. Comparative analyses highlight the dynamic nature of tunicate genomes, with conserved macrosynteny yet extensive microsyntenic rearrangements and scrambling, underscoring their rapid evolutionary trajectory. This high-quality genome assembly provides a valuable resource for exploring the unique biological features of colonial chordates, including their exceptional regenerative abilities and complex allorecognition system.

Animals

Annotated genome assemblies of two temperate North American dung beetles, Canthon chalcites and Phanaeus vindex.

Dung beetles serve as cultivators of their natural habitats, improving soil health and functions in both natural and anthropogenic environments. Despite their ecological importance, whole genome sequences for Scarabaeinae are limited. Here, we present the draft annotated genome assemblies for 2 temperate species of North American dung beetles collected from eastern Tennessee: Canthon chalcites and Phanaeus vindex. Both genome assemblies were generated from PacBio long reads and have high completeness, with BUSCO scores of 98.1% and 98.6% for C. chalcites and P. vindex, respectively. For C. chalcites, the BRAKER3 pipeline predicted 12,799 genes, and the gene set was 93.7% complete. For P. vindex, the BRAKER3 predicted 12,252 genes, and the gene set was 94.9% complete. From the annotated gene sets, orthologous protein sequence analyses among C. chalcites, P. vindex, the dung beetle species Onthophagus taurus, and the more evolutionarily distant beetle Tribolium castaneum indicated that there are 260 unique protein clusters for C. chalcites and 210 unique protein clusters for P. vindex. These 2 draft genomes provide valuable data for comparative genomics, evolution, and phylogenic studies for dung beetle species.

Animals

Chromosome-Level Genome Assembly and Annotation of the Chinese Lizard Gudgeon (Saurogobio dabryi).

The Chinese lizard gudgeon (Saurogobio dabryi) is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in China, the lack of genomic resources has rendered the genetic breeding and conservation research. Here, we present the first chromosome-level genome assembly of S. dabryi using PacBio HiFi long reads, short reads, and Hi-C sequencing data. The final assembly reaches a total size of 1.09 Gb and Hi-C scaffolding anchors 99.55% of the assembled contigs onto 25 chromosomes, with a scaffold N50 reaching 43.15 Mb. The final genome assembly shows a BUSCO completeness of 98.39%. We annotated 659.55 Mb repetitive sequences and 26,036 protein-coding genes, 99.47% of which are functionally annotated. Comparative phylogenomic analysis clarifies the phylogenetic position of Saurogobio within Gobioninae. This high-quality genome provides a critical genetic basis for exploring cyprinid phylogeny, benthic adaptive evolution, genetic improvement, and conservation efforts of S. dabryi.

Saurogobio dabryi

De novo genome assemblies of threatened Asian hornbills (Bucerotidae) reveal declining population trajectories during the late Pleistocene.

BACKGROUND: Asian hornbills are flagship species of the wet tropics that face significant threats from hunting, habitat loss, and fragmentation. Despite being conservation flagships, whole genome information is available for only two of the 32 Asian hornbill species. In this study, we provide the first de novo genome assemblies for four hornbill species (Bucerotidae) in Asia. METHODS: We used a combination of long-read and short-read sequencing data to assemble and annotate de novo hybrid genomes of four species of hornbills. We also assembled and compared mitochondrial genomes of these species. Using a comparative genomics approach, we performed orthology assignment and gene evolution analyses to identify unique gene families in Asian hornbills, gene families that showed significant expansion, their functions and structural variation. Furthermore, using the Pairwise Sequentially Markov Coalescent (PSMC) method, we reconstructed demographic histories of hornbill species to examine changes in their population trajectories in the past. RESULTS: We present hybrid genome assemblies for Great Hornbill (B. bicornis - GH), Rufous-necked Hornbill (A. nipalensis- RNH), Malabar Pied Hornbill (A. coronatus- MPH) and Wreathed Hornbill (R. undulatus- WH). The genome sizes of these hornbills range from 1.1 Gb to 1.3 Gb, with over 95.9% completeness and gene prediction BUSCO. We reported 10,525 orthogroups shared among four Asian hornbill species and identified significant expansion in gene families associated with structural keratin development in Asian hornbills compared to their ancestors. We also provide annotated mitogenomes for each of these species. Furthermore, we found that the WH, a more abundant, widely distributed, and migratory species, showed a higher Ne than the other three hornbill species. However, an overall decline in Ne for all species was recorded during the Pleistocene climatic fluctuations. CONCLUSIONS: We present the first-ever, high-quality reference genomes for the threatened hornbill species from Asia. Hornbills have shown significant expansion in genes involved in structural keratin development. Our results indicate that Pleistocene climatic fluctuations have led to dramatic population declines in all four species. We believe that this study provides robust genomic resources to support future comparative and conservation genomics efforts for hornbills.

Animals

Chromosome-level genome assembly of Elaeocarpus petiolatus (Elaeocarpaceae).

Elaeocarpus petiolatus is an ecologically and economically important species in tropical and subtropical forests. Despite its significance, the lack of genomic resources has hindered research on the genetic diversity and adaptive traits of E. petiolatus. To address this gap, we present a comprehensive chromosome-level genome assembly of E. petiolatus generated using advanced PacBio high-fidelity (HiFi) long-read sequencing and Hi-C technology. The assembly spans 322.45&#x2009;Mb, with a scaffold N50 of 20.58&#x2009;Mb, indicating that 37.11% of the genome is composed of repetitive elements. We identified 25,295 protein-coding genes, of which 96.74% were functionally annotated. This high-quality genome provides a critical resource for understanding the genetic mechanisms underlying environmental adaptability and biosynthesis of bioactive compounds in E. petiolatus, thereby supporting conservation efforts and sustainable forest management. The assembled genome and associated sequencing data are publicly available, facilitating further evolutionary and functional studies on the Elaeocarpaceae family.

Chromosomes, Plant

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis

Complete telomere-to-telomere genome assembly of Guazuma ulmifolia uncovers evolutionary mechanisms, drought adaptation, and flavonoid biosynthesis.

The first T2T reference genome of Guazuma ulmifolia is reported, which serves as a core genomic resource for stress adaptation research and stress-tolerant breeding in cacao wild relatives. Climate change, particularly increased incidence of drought, poses a major threat to food security. Understanding the genomic basis of environmental adaptation in crop wild relatives can provide valuable resources for improving stress resilience. Guazuma ulmifolia, a wild relative of Theobroma cacao with important ecological and medicinal value, lacks high-quality reference genomic resources. Here, we report the first telomere-to-telomere (T2T) chromosome-level genome assembly of G. ulmifolia, with a genome size of 311.31&#xa0;Mb, contig N50 of 35.19&#xa0;Mb, and 98.70% BUSCO completeness. Repetitive sequences constitute 27.43% of the G. ulmifolia genome, with LTR retrotransposons as the predominant class. Comparative genomic analyses revealed that genome-size variation among Malvaceae species is associated with differences in polyploidization history and TE dynamics. Ancestral karyotype reconstruction identified five lineage-specific chromosome fusion events distinguishing G. ulmifolia from T. cacao. Comparative analyses further identified tandem duplication-associated expansion of stress-related LEA and GST gene families, suggesting potential genomic features associated with stress responses. Flavonoid biosynthesis genes were largely conserved in copy number but showed tissue-specific expression patterns, providing candidate genes for investigating secondary metabolism. Together, this study establishes a high-quality T2T genome resource for exploring genome evolution, chromosome organization, and stress-related genomic features in Malvaceae.

Genome, Plant

Chromosome-level genome assembly of the large carpenter bee Xylocopa dejeanii Lepeletier, 1841 (Hymenoptera: Apidae).

Xylocopinae, a diverse bee subfamily comprising over 1,000 bee species, and also a major model system for studying the pollination and evolution of sociality. The lack of chromosome-level genome assembly resources for the Xylocopinae limits our research of their biology and evolution. Here, we provided the first pseudo-chromosomes genome assembly of the Xylocopa dejeanii combined PacBio CLR long reads, Illumina sequences, and Hi-C data. The final genome is 194.44&#x2009;Mb located in 16 chromosomes. Our assembly includes 141 scaffolds, with a scaffold N50 length of 13.15&#x2009;Mb. BUSCO analysis revealed 99.00% completeness. Genome annotation identified 28.27&#x2009;Mb of repetitive elements, 10,970 protein-coding genes, and 432 ncRNAs. This high-quality X. dejeanii assembly advances our understanding of Xylocopinae genomics and provides new insights into bee evolution.

Animals