Search PubMedSearch

SEARCH · Search PubMed

Results for “de novo assembly”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

De novo assembly and authentication of ancient DNA metagenomes with nf-core/mag.

Ancient DNA provides a direct window into the evolutionary processes that have shaped living microbial species today, as well as their now extinct relatives. Advances in both sequencing methods and de novo assembly techniques have not only resulted in a flood of modern metagenomic sequencing data, but they have also allowed palaeogenomicists to retrieve vast amounts of ancient DNA from past microorganisms, including species and strains without modern reference genomes. However, the degraded nature of ancient DNA means that the standard techniques of genome assembly developed for modern DNA are unlikely to perform effectively, unless heavily modified. This hinders the incorporation of ancient data into broader metagenomic studies that would otherwise benefit from having deep time information on the evolution of different microbial species. In this primer and protocol paper, we provide guidance on ways to adapt existing metagenomic de novo assembly processes, including data input, tools, and settings, in order to perform more robustly and effectively on ancient DNA. After assembly, we then further describe how ancient DNA contigs can be identified and validated. The key steps of ancient metagenomic assembly are now integrated in a dedicated ancient DNA mode in the established pipeline nf-core/mag. By introducing support for ancient DNA data in nf-core/mag, we aim to improve the ability of researchers to more regularly integrate de novo assembled ancient microbial data into broader metagenomics studies of microbial ecology and evolution.

DNA, Ancient

De novo assembly of transcriptomes of six Hua species (Semisulcospiridae, Cerithioidea, Gastropoda).

Species in Semisulcospiridae are important in freshwater ecology and have great research value, yet their genomic resources remain very limited. Here, we present de novo assembled transcriptomes from six species of Hua in Semisulcospiridae, including Hua textrix (Heude, 1888), H. yangi L.-N. Du, J.-X. Yang & Chen, 2023, H. wujiangensis L.-N. Du, J.-X. Yang & Chen, 2023, and three undescribed species. Assembly was performed using Trinity, resulting in average contig lengths ranging from 716.6 to 883.3 bp and transcript numbers ranging from 147,147 to 268,741. Benchmarking Universal Single-Copy Ortholog (BUSCO) analysis was used to assess the transcriptome completeness. The functional annotation of transcripts for each species had over 18,000 BLAST hits, 17,000 GO terms, 15,000 KEGG pathways, 8,000 Pfam accessions, and 140 COG functional categories. This study provides valuable transcriptomic resources for the six Hua species, which can be used for various research of Semisulcospiridae, including biodiversity, phylogeny, and comparative genomics.

Transcriptome

De Novo Assembly and Comparative Analysis of the Complete Mitochondrial Genome of Mesenchytraeus (Annelida, Enchytraeidae).

The Changbai Mountain range is one of the key glacial refugia in Northeast Asia. Mesenchytraeus exhibits high species diversity, strong endemism, and widespread cryptic species in this region, for which mitogenomes provide useful molecular markers for exploring cryptic species complexes. This makes Mesenchytraeus an ideal model for studying mitogenome evolution among closely related lineages; however, no mitogenome data have been reported for this genus to date. In this study, we performed de novo assembly, annotation, and comparative analysis of the mitogenomes of 13 Mesenchytraeus species (14 individuals) from Changbai Mountain. All mitogenomes are typical circular molecules containing 37 genes, but putative control regions are rearranged and consistently located between ATP6 and trnR. All species exhibit annelid-specific strand nucleotide biases, characterized by negative GC skew and near-zero AT skew. Codon usage analysis reveals that codon families with wobble U are significantly biased toward mtDNA codons, whereas those with wobble C or G are biased toward non-mtDNA codons, suggesting a conserved mitochondrial codon usage pattern in annelids. All tRNAs form typical cloverleaf secondary structures except trnS2, which lacks the D-stem and the dihydrouridine (DHU) arm in some species. The putative control regions commonly contain complex palindromic repeats, hairpins, and repetitive elements, and may harbor dual replication origins. Phylogenetic analyses support the monophyly of Mesenchytraeus and reveal significant molecular divergence among morphologically cryptic species. This study provides the first mitogenome dataset for Mesenchytraeus and offers new insights into the evolution and replication mechanisms of mitogenomes in Clitellata and broader Annelida.

Mesenchytraeus

Colora: a Snakemake workflow for complete chromosome-scale de novo genome assembly.

MOTIVATION: De novo assembly creates reference genomes that underpin many modern biodiversity and conservation studies. Large numbers of new genomes are being assembled by labs around the world. To avoid duplication of efforts and variable data quality, we desire a best-practice assembly process, implemented as an automated portable workflow. RESULTS: Here, we present Colora, a Snakemake workflow that produces chromosome-scale de novo primary or phased genome assemblies complete with organelles using Pacific Biosciences HiFi, Hi-C, and optionally Oxford Nanopore Technologies reads as input. Colora is a user-friendly, versatile, and reproducible pipeline that is ready to use by researchers looking for an automated way to obtain high-quality de novo genome assemblies. AVAILABILITY AND IMPLEMENTATION: The source code of Colora is available on GitHub (https://github.com/LiaOb21/colora) and has been deposited in Zenodo under DOI https://doi.org/10.5281/zenodo.13321576. Colora is also available at the Snakemake Workflow Catalog (https://snakemake.github.io/snakemake-workflow-catalog/? usage=LiaOb21%2Fcolora).

Software

Chromosome-level de novo assembly of the nuclear and mitochondrial genomes of Arcopilus aureus, a filamentous fungus with multifaceted ecological and economic roles.

The filamentous fungus Arcopilus aureus (Sordariale: Chaetomiaceae) is notable for its multi-domain significance across agriculture, medicine, and industry. In this study, we generated a chromosome-level nuclear genome and a complete circular mitogenome for A. aureus by integrating data from next-generation sequencing, PacBio HiFi, and Hi-C technologies. The final nuclear genome assembly spans 33.77 Mb (GC content: 57.67%), and was organized into seven chromosomal-sized scaffolds (only one gap) with an N50 size of 5.09 Mb and BUSCO completeness of 95.91%. A total of 10,282 protein-coding genes, 228 non-coding RNAs, and ~1.77 Mb of repetitive elements were predicted in the nuclear genome. By contrast, the mitogenome of A. aureus is 33,820 bp in length, with a GC content of 25.96%. It harbors 15 typical mitochondrial protein-coding genes, one unidentified ORF, two rRNAs (small subunit rns and large subunit rnl), and 28 tRNAs. This high-quality genome assembly provides a valuable resource for understanding the ecology, genetics, and evolution of A. aureus, which facilitates elucidating its mechanisms of biocontrol, infection, and metabolite synthesis.

Genome, Mitochondrial

De Novo Assembly of the Trypanosoma congolense Genome Reveals an Organization Influenced by Antigenic Variation but Distinct from Trypanosoma brucei.

Antigenic variation allows pathogens to evade mammalian adaptive immunity through the continuous change in exposed antigens. In African trypanosomes, antigenic variation involves changes in expressed Variant Surface Glycoproteins (VSGs). Understanding of VSG expression control and change amongst African trypanosomes is most advanced in Trypanosoma brucei. In the important animal trypanosome, Trypanosoma congolense, incomplete genome assembly has held back understanding of the mechanics of antigenic variation. Here, we have used long-read DNA sequencing and Hi-C DNA interaction analysis to provide a telomere-to-telomere assembly of the T. congolense genome. This assembly reveals a genome comprising 12 diploid chromosomes, one tetraploid chromosome, and more than 100 small chromosomes. With this assembly we reveal several features of VSG organization and expression that differ from T. brucei. The majority of the T. congolense VSG archive, estimated at ∼1,500 genes, localizes to subtelomeres in 12 of the 13 large chromosomes, but these loci are notably smaller than are found in T. brucei. Furthermore, transcriptome analysis suggests expression of VSGs across the T. congolense subtelomeres, which are not separated within the nucleus from non-VSG chromosome regions, suggesting that there is no dedicated VSG expression site. Strikingly, one chromosome contains approximately 40% of the VSG archive and is largely transcriptionally silent, potentially acting as the major reservoir of new VSG variants. Finally, we show that VSG expression can be detected from multiple small chromosomes. In summary, the new genome assembly provides a platform for understanding a potentially unusual operation of VSG expression and switching in T. congolense.

Trypanosoma congolense

De novo genome assemblies of threatened Asian hornbills (Bucerotidae) reveal declining population trajectories during the late Pleistocene.

BACKGROUND: Asian hornbills are flagship species of the wet tropics that face significant threats from hunting, habitat loss, and fragmentation. Despite being conservation flagships, whole genome information is available for only two of the 32 Asian hornbill species. In this study, we provide the first de novo genome assemblies for four hornbill species (Bucerotidae) in Asia. METHODS: We used a combination of long-read and short-read sequencing data to assemble and annotate de novo hybrid genomes of four species of hornbills. We also assembled and compared mitochondrial genomes of these species. Using a comparative genomics approach, we performed orthology assignment and gene evolution analyses to identify unique gene families in Asian hornbills, gene families that showed significant expansion, their functions and structural variation. Furthermore, using the Pairwise Sequentially Markov Coalescent (PSMC) method, we reconstructed demographic histories of hornbill species to examine changes in their population trajectories in the past. RESULTS: We present hybrid genome assemblies for Great Hornbill (B. bicornis - GH), Rufous-necked Hornbill (A. nipalensis- RNH), Malabar Pied Hornbill (A. coronatus- MPH) and Wreathed Hornbill (R. undulatus- WH). The genome sizes of these hornbills range from 1.1 Gb to 1.3 Gb, with over 95.9% completeness and gene prediction BUSCO. We reported 10,525 orthogroups shared among four Asian hornbill species and identified significant expansion in gene families associated with structural keratin development in Asian hornbills compared to their ancestors. We also provide annotated mitogenomes for each of these species. Furthermore, we found that the WH, a more abundant, widely distributed, and migratory species, showed a higher Ne than the other three hornbill species. However, an overall decline in Ne for all species was recorded during the Pleistocene climatic fluctuations. CONCLUSIONS: We present the first-ever, high-quality reference genomes for the threatened hornbill species from Asia. Hornbills have shown significant expansion in genes involved in structural keratin development. Our results indicate that Pleistocene climatic fluctuations have led to dramatic population declines in all four species. We believe that this study provides robust genomic resources to support future comparative and conservation genomics efforts for hornbills.

Animals

Ulmus minor response to Dutch elm disease: de novo transcriptome assembly and annotation.

Dutch elm disease (DED), caused by Ophiostoma novo-ulmi (ONU), has devastated elm populations across Europe and North America since the 20th century. In this work, a de novo transcriptome assembly of Ulmus minor in response to ONU is presented. We used two DED-resistant genotypes, MDV2.3 and VAD2, and one DED-susceptible genotype, MDV1, to capture responses to ONU at four time points post-inoculation (6, 24, 72, and 144 hours). RNA from collected samples was isolated and sequenced producing 60.88 M 100 bp paired-end reads per sample. We performed a de novo transcriptome assembly combining data from the three genotypes. The assembly was functionally annotated and validated through differential gene expression analysis of the response. This dataset provides a valuable resource for studying molecular mechanisms of DED resistance in elms, contributing to broadening our understanding of tree immunity and facilitating potential applications in functional annotation of future genome assemblies.

Transcriptome

De novo genome assembly of Clonostachys rosea CMAA1284: a fungal strain with biotechnological potential.

Clonostachys rosea is a fungus with significant applications in biological control, plant growth promotion, and secondary metabolite production essential for agriculture. We report the de novo genome assembly of C. rosea strain CMAA1284 (also known as LQC 62), which shows high contiguity and excellent gene completeness, supporting its use for functional genomics and biotechnological studies.

Clonostachys rosea

De novo transcriptome assembly and gene expression analysis of Cnidium officinale under high-temperature conditions.

BACKGROUND: The medicinal plant Cnidium officinale (CO) is widespread in Northeast Asia and vulnerable to heat stress. The naturally occurring composition of pharmacological ingredients of CO results in overall physiological consequences; therefore, it is crucial to have a comprehensive understanding of metabolic response to ambient heat in terms of acclimation to estimate how much CO is exposed to threatening environmental conditions. RESULTS: Transcriptome analysis is critical for understanding the consequences of long-term physiological adaptation of CO to abiotic stress. However, transcriptome analysis on this species, particularly under prolonged stress conditions, has remained limited. We employed a temperature gradient tunnel (TGT) to subject CO to high-temperature exposure for four months, enabling us to observe the cumulative effects of heat and assess its acclimation mechanisms. In the absence of genome sequencing data, we performed de novo transcriptome assembly and compared DEGs from temperature treatment plots of a TGT and a growth chamber (GC). Since interpreting transcriptomic data can be complex, we employed a sequential analytical approach, including DEG clustering, GO enrichment, KEGG pathway mapping, miRNA-target gene analysis, and multiple rounds of RNA sequencing validation. DEGs were classified into two categories: genes exhibiting significant fold changes and genes showing significant count changes rather than fold changes. Then, we analyzed the functional roles of DEGs to determine which pathways respond to ambient and stressful high temperatures and validated the findings through cross-comparison with GC. Additionally, we conducted miRNA analysis to investigate post-transcriptional regulation under high temperatures. CO grown under higher ambient temperatures exhibited slight upregulation of pathways related to protein stability and turnover, ABA biosynthesis, and energy production, such as photosynthesis and oxidative phosphorylation. However, under extreme heat stress, most metabolic pathways were downregulated except for those involved in transcription, translation, oxidative phosphorylation and the biosynthesis of cutin, suberin, and wax. CONCLUSION: This study demonstrated that proper clustering of genes based on expression levels and fold changes in two different experimental conditions, along with pathway mapping, may provide a comprehensive understanding of CO's response to heat stress. These insights could contribute to future research on heat tolerance and crop improvement.

Gene Expression Profiling

Genome Report: De novo genome assembly of the greater Bermuda land snail, Poecilozonites bermudensis (Mollusca: Gastropoda), confirms ancestral genome duplication.

Poecilozonites bermudensis, the greater Bermuda land snail, is a critically endangered species and one of only two extant members in its genus. These snails are one of Bermuda's few endemic animal clades and their rich fossil record was the basis for the punctuated equilibria model of speciation. Once thought extinct, recent conservation efforts have focused on the recovery of the species, yet no genomic information or other molecular sequences have been available to inform these initiatives. We present a high-quality, annotated genome for P. bermudensis generated using PacBio long read and Omni-C short read sequencing. The resulting assembly is approximately 1.36 Gb with a scaffold N50 of 44.t Mb and 31 chromosome-length scaffolds. Nearly 43 percent of the genome was identified as repeat content. This assembly will serve as a resource for the conservation and study of P. bermudensis, and its only close extant and also critically endangered relative, P. circumfirmatus. Additionally, this genome adds to the growing body of data needed for a more complete understanding of gastropod evolution and for evolutionary processes in general.

Annotation

Rawsamble: overlapping raw nanopore signals using a hash-based seeding mechanism.

MOTIVATION: Raw nanopore signal analysis is a common approach in genomics to provide fast and resource-efficient analysis without translating the signals to bases (i.e. without basecalling). However, existing solutions cannot interpret raw signals directly if a reference genome is unknown due to a lack of accurate mechanisms to handle increased noise in pairwise raw signal comparison. Our goal is to enable the direct analysis of raw signals without a reference genome. To this end, we propose Rawsamble, the first mechanism that can identify regions of similarity between all raw signal pairs, known as all-vs-all overlapping, using a hash-based search mechanism. RESULTS: We use these overlaps to construct de novo assembly graphs with an existing assembler, miniasm, off-the-shelf. To our knowledge, these are the first de novo assemblies ever constructed directly from raw signals without basecalling. Our extensive evaluations across multiple genomes of varying sizes show that Rawsamble provides a significant speedup (on average by 5.01× and up to 23.10×) and reduces peak memory usage (on average by 5.74× and up to by 22.00×) compared to a conventional genome assembly pipeline using the state-of-the-art tools for basecalling (Dorado's fastest mode) and overlapping (minimap2) on a CPU. We find that around one-third of Rawsamble's overlapping pairs are also found by minimap2. We find that when we use overlapping reads from Rawsamble, we can construct unitigs that are (i) as accurate as those built from minimap2's overlaps and (ii) up to half a chromosome in length (e.g. 2.3 million bases for E. coli). AVAILABILITY AND IMPLEMENTATION: Rawsamble is available at https://github.com/CMU-SAFARI/RawHash. We also provide the scripts to fully reproduce our results on our GitHub page.

Nanopores

Detection and characterization of neonatal cytomegalovirus through nanopore sequencing using flongle flow cells: Pilot study in Philadelphia, Pennsylvania.

BACKGROUND: Cytomegalovirus (CMV) remains a significant infection in neonates and its early detection can aid with further treatment (antiviral, audiology). However, current diagnostics do not provide genetic information. OBJECTIVE: We explored the use of the portable and comprehensive sequencing method from Oxford Nanopore Technologies, utilizing low-cost Flongle flow cells to detect and perform sequence-level characterization of neonatal urine samples that tested positive for CMV by PCR. STUDY DESIGN: We performed a pilot study based on a retrospective cohort study of neonates who were positive for CMV by PCR, who were admitted at two birth hospitals in Philadelphia, PA. We leveraged deep and long-read sequencing results to analyze the reads in two forms: by comparing them against a reference-based strain and by reconstructing the genome through de novo assembly with phylogenetic tree analysis. RESULTS: We assayed seven clinical samples, including a positive and negative control sample, from newborns ranging from 23 weeks' gestation to term, with testing performed for microcephaly, hearing test results, small gestational age, and thrombocytopenia. Each sample showed multiple differences compared to the reference strain, and the phylogenetic tree analysis of the de novo assembly depicted the genetic diversity of the samples. CONCLUSION: This pilot study shows that nanopore sequencing with low-cost Flongle flow cells can detect and characterize CMV strains from clinical neonatal urine samples. This, coupled with current screening and diagnostic criteria, could further our genomic understanding of neonatal CMV, such as viral genome diversity, genotype-phenotype associations, and spread of strains.

Humans

Metatranscriptomic analysis of viral sequences associated with Culex nigripalpus at an Alabama aquaculture site.

Mosquitoes associated with aquaculture habitats can harbor diverse viruses, yet the viromes of many locally abundant species remain poorly characterized. At an aquaculture-associated site in Auburn, Alabama, we surveyed mosquito populations and found Culex nigripalpus to be the dominant species collected. To characterize viruses associated with this mosquito, we performed RNA-seq on pooled female Cx. nigripalpus and compared complementary bioinformatic workflows for viral detection and genome recovery. One workflow removed host-associated reads by mapping to the closest available mosquito reference genome prior to assembly, whereas a second workflow used fully de novo assembly and viral database annotation. Additional protein-level filtering, cross-workflow comparison, and comparison of Trinity and rnaSPAdes assemblies were used to prioritize well-supported viral candidates. Across the original analyses, 16 submitted accessions corresponding to 12 collapsed virus/name groups were recovered, including Merida virus, Hubei mosquito virus 5, Zhejiang mosquito virus, Hubei virga-like virus 3, Rinkaby virus, Elemess virus, Qingnian mosquito virus, Serbia narna-like virus 2, XiangYun narna-levi-like virus 8, Ecclesville picorna-like virus, and baculovirus-like fragments. Several candidates were supported across multiple workflows, while others were recovered only under specific analytical conditions, indicating that candidate recovery was influenced by assembly and filtering choices. Selected viral contigs were independently supported by RT-PCR amplification. Overall, these results provide a first characterization of viral sequences associated with Cx. nigripalpus from an Alabama aquaculture-associated site and show that comparison across assembly and filtering strategies helped prioritize the most consistently supported viral candidates.

Animals

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article

High quality genome assemblies of African cattle breeds using PacBio HiFi sequencing.

Africa has a uniquely rich cattle diversity of ~150 breeds comprising the Bos taurus indicus sub-species, Bos taurus taurus, and their crosses. These represent ~23% of the global cattle population. However, high quality, representative assemblies are limited for African cattle and especially for indicine breeds. Here we built high quality de novo assemblies for five important African indigenous cattle breeds using PacBio HiFi sequencing: Lagune (Bos taurus taurus), Gudali, Iringa Red and Singida White (Bos taurus indicus), and Mpwapwa (Bos taurus taurus x Bos taurus indicus). These new assemblies are the most contiguous and complete African cattle assemblies produced so far, with genome sizes of 3.25-3.36 Gb, contiguity N50s ranging from 83.59 Mb to 97.87 Mb and scaffold N50s from 100.30 Mb to 113.37 Mb. BUSCO genome completeness scores were also higher than 99.68%, indicative of highly contiguous assemblies. These improved and highly contiguous genome assemblies are consequently a valuable resource for future African and global livestock genomic studies.

Animals

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article