Search PubMedSearch

PubMed · 42595963

Plant species identification by genome skimming across the vascular plant tree of life.

Abstract

Accurate species identification is essential for biodiversity conservation and sustainable use, yet standard plant DNA barcoding often fails to achieve species-level resolution. We present a large-scale empirical evaluation of genome skimming as a tool to improve plant species discrimination. Using standardised data from 1969 individuals representing 475 species from 32 genera across major lineages of the vascular plant tree of life, we compare conventional plastid + internal transcribed spacer (ITS) barcodes with genome skimming approaches. Standard barcoding using rbcL, matK, trnH-psbA and ITS resolved about half of species (49.3%), with six genera showing <&#x2009;25% species discrimination. By contrast, genome skimming enabled the recovery of complete plastid genomes, yielding 57.6% species discrimination. It also generated sufficient nuclear genomic data for additional resolution from k-mer analysis, achieving 66.8% species discrimination - an average gain of 17.5% over standard barcodes - while eliminating cases of extreme failure (<&#x2009;25% resolution). The recovery of complete plastomes and ribosomal DNAs from genome skims also ensures backward compatibility with existing barcode datasets. Our results demonstrate that genome skimming provides data that substantially improves species-level resolution across diverse plant lineages and offers a scalable, high-throughput approach for building comprehensive reference resources to support global biodiversity initiatives.

Explore related subjects

Keep this discovery

BibTeXRIS

Chun-Xia Zeng, Meng-Yuan Zhou, Alex D Twyford, Lian-Ming Gao, Zhi-Yong Zhang, Yun-Heng Ji, Hong-Tao Li, Peng-Fei Ma, Wen-Bin Yu, Zhi-Qiong Mo, Zheng-Yu Zuo, Wu Huang, Ting-Shuang Yi, Jun-Bo Yang, Peter M Hollingsworth, De-Zhu Li. 2026-08-13. Plant species identification by genome skimming across the vascular plant tree of life.. https://doi.org/10.1111/nph.71494

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related citations

Performance comparison of rapid and native barcoding methods for Oxford Nanopore sequencing of Poliovirus Viral Protein 1 (VP1) amplicons.

Accurate and timely sequencing of poliovirus is critical for global eradication efforts, particularly for molecular epidemiology based on the typing region of the genome, viral protein 1 (VP1). While Oxford Nanopore Technologies (ONT) sequencing has expanded capabilities for poliovirus surveillance, the relative performance of different ONT library preparation methods, including ligation-based (Native Barcoding) and transposase-based (Rapid Barcoding) approaches, has not been systematically evaluated. In this study, we compared rapid barcoding and native barcoding workflows for sequencing VP1 amplicons from 17 type 2 poliovirus-positive samples, each processed in triplicate. Native barcoding generated significantly more sequencing output, producing approximately 2.3-fold greater total read yield than rapid barcoding, and demonstrated higher run-to-run reproducibility (R2 = 0.979-0.998 vs. 0.847-0.929, respectively; p&#x202f;<&#x202f;0.001). In addition, native barcoding generated 80% of the total yield achieved by rapid barcoding within approximately 7&#x202f;h, whereas rapid barcoding required approximately 40&#x202f;h to reach the same output. Despite these differences, both methods produced identical VP1 consensus sequences across all samples, with comparable read quality (median per-base Q-scores of approximately Q17-Q18). Rapid barcoding provided substantial practical advantages, reducing hands-on library preparation time (55 vs. 200&#x202f;min) and per-sample cost ($12.82 vs. $16.54), while simplifying workflow and reducing technical complexity. These findings indicate that sequencing yield may not be a determinant of downstream analytical outcomes for poliovirus VP1 ONT sequencing. Rapid barcoding therefore represents a cost-effective and efficient approach for routine poliovirus surveillance, whereas native barcoding remains advantageous in applications requiring rapid data generation or maximal sequencing depth.

Poliovirus

Getting to the Core of the Matter-Assessing the Role of Replication in Metabarcoding-Based sedaDNA.

Replication is central to most experimental and sampling designs, increasing inferential power and capturing fine-scale data heterogeneity. However, its importance remains poorly evaluated in some ecological and evolutionary settings. This is the case of metabarcoding studies using DNA recovered from sedimentary archives, in which biological signals integrate ecological information through depositional and burial processes, yet are commonly inferred from a single sediment core per site. Here, we evaluated the effect of different types of replication using sedimentary DNA metabarcoding data from two genetic markers (mitochondrial COI and nuclear 18S) using a nested sampling design. The design included three intertidal sites, three spatially separated sediment cores per site (biological replicates), two sediment horizons per core, and eight PCR (technical) replicates per sediment sample. Variance partitioning showed that site identity and sediment age group together explained >&#x2009;70% of the variation in beta diversity, indicating that among-site spatial and stratigraphic differences were the dominant drivers of community composition. PERMANOVA likewise identified non-significant effects of biological replication. Among PCR replicates from the same sediment sample, richness varied substantially, whereas Shannon diversity was more consistent. Despite this variability, differences in community composition among technical replicates remained smaller than those associated with biological replication or site identity, indicating a limited influence on broader ecological patterns. Community composition was highly similar among replicate cores within sites, consistent with stratigraphic coherence. These results indicate limited within-site heterogeneity and suggest that, under stratigraphically coherent conditions, increasing biological replication may provide little additional information, whereas enhancing technical replication and stratigraphic resolution can improve ecological inference from sedimentary DNA metabarcoding datasets.

DNA Barcoding, Taxonomic

Cytonuclear conflict and reticulate evolution in the Morelloid clade (Solanum, Solanaceae): Insights from genome skimming and network Phylogenomics.

The Morelloid clade (black nightshades) is one of the most strongly supported clades within the megadiverse Solanum genus. It comprises 76 globally distributed, non-spiny herbaceous and suffrutescent species. While often erroneously considered poisonous weeds, several species are economically important as orphan crops. The clade is closely related to tomato and potato but, due to a lack of focused breeding efforts, remains a putative reservoir of genetic diversity for crop improvement. Despite this potential, we lack fundamental knowledge on the evolution of the Morelloid clade. The group includes polyploid species with unknown parental origins-likely reflecting reticulate processes such as hybridization, introgression, and associated backcrossing events. Prior analyses have been unable to disentangle these processes, leaving the mechanisms underlying reticulate evolution in the Morelloid clade poorly understood. Here, we use genome skimming to produce a well-supported maximum likelihood plastid phylogeny from complete circularized plastomes and a coalescent-based species tree from combined Angiosperms353 and conserved ortholog set nuclear markers. Our dataset, composed of previously published data and deep genome skimming from herbarium samples, spans 26 Morelloid species. To investigate phylogenetic discordance, we used a nuclear phylogenetic network, multispecies coalescent simulations, a fused rooted nuclear chloroplast tree, and quantification of nuclear gene tree concordance. We show that incongruence between nuclear and plastid trees is pervasive and cannot be explained by incomplete lineage sorting alone. Instead, our results demonstrate that events consistent with repeated chloroplast capture have shaped the reticulate evolutionary history of the clade, especially among African polyploid and Pan-American diploid lineages.

Phylogeny