Search PubMedSearch

SEARCH · Search PubMed

Results for “capture enrichment sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Culture-free genomics: a shift toward genome-wide applications in Chagas disease and leishmaniasis.

INTRODUCTION: Chagas disease and leishmaniasis remain major neglected tropical diseases, with diagnosis and surveillance constrained by low parasite burden, multiclonal infections, and complex parasite biology. Traditional culture-dependent and targeted molecular approaches fail to capture the full genomic diversity of Trypanosoma cruzi and Leishmania spp. limiting clinical and epidemiological utility. AREAS COVERED: We review the evolution from early sequencing to second- and third-generation platforms, highlighting culture-free detection and genomic surveillance. We discuss enrichment strategies (selective whole-genome amplification (SWGA) and capture-enrichment sequencing (CES)) addressing low parasite DNA abundance in complex samples, alongside metagenomics and portable sequencing for field-based surveillance and diagnostics. We further explore how direct-from-host data can improve diagnostics, enhance transmission surveillance, support treatment monitoring, and guide control strategies. EXPERT OPINION: Culture-free genomic approaches represent a transformative advance in kinetoplastid research, providing resolution that culture-dependent methods cannot deliver. Their diagnostic contribution is at present largely indirect, operating through the identification of improved molecular and serological targets rather than through sequencing as the assay itself. Persistent barriers of cost, infrastructure, standardization, and bioinformatics capacity, together with the absence of formal clinical validation, currently confine these methods to research and surveillance settings.

Capture-enrichment sequencing

Phylogenomic reconstruction of Cryptosporidium spp. captured directly from clinical samples reveals extensive genetic diversity.

Cryptosporidium is a leading cause of severe diarrhea and mortality in young children and infants in Africa and southern Asia. More than twenty Cryptosporidium species infect humans, of which C. parvum and C. hominis are the major agents causing moderate to severe diarrhea. Relatively few genetic markers are typically applied to genotype and/or diagnose Cryptosporidium. Most infections produce limited oocysts making it difficult to perform whole genome sequencing (WGS) directly from stool samples. Hence, there is an immediate need to apply WGS strategies to 1) develop high-resolution genetic markers to genotype these parasites more precisely, 2) to investigate endemic regions and detect the prevalence of different genotypes, and the role of mixed infections in generating genetic diversity, and 3) to investigate zoonotic transmission and evolution. To understand Cryptosporidium global population genetic structure, we applied Capture Enrichment Sequencing (CES-Seq) using 74,973 RNA-based 120 nucleotide baits that cover ~92% of the genome of C. parvum. CES-Seq is sensitive and successfully sequenced Cryptosporidium genomic DNA diluted up to 0.005% in human stool DNA. It also resolved mixed strain infections and captured new species of Cryptosporidium directly from clinical/field samples to promote genome-wide phylogenomic analyses and prospective GWAS studies.

Cryptosporidium

Mannheimia haemolytica strain-level diversity in cattle populations.

High-resolution genomic characterization is essential for understanding diversity, pathogenicity, and transmission dynamics of bacterial pathogens. Mannheimia haemolytica (Mh) is the most consequential bacterial agent associated with bovine respiratory disease (BRD) in cattle, as a leading cause of morbidity, mortality, and antimicrobial use. Historically, BRD pathogens, including Mh, have been studied using culture or PCR approaches that provided limited ability to characterize fine-scale genomic variation across communities. Here, we evaluated target-enriched (TE) shotgun sequencing, a culture-independent method capable of strain-level resolution within metagenomic data, for detecting and characterizing Mh in comparison with qPCR and 16S rRNA gene sequencing. Nasal swabs (10 individual and 2 composited DNA samples per pen) and environmental samples (three ropes hung on pen rails and three water bowl swabs per pen) were collected from four pens in each of five distinct cattle populations. DNA was extracted for TE sequencing to identify Mh at both species and genomic sequence variant (GSV) levels, and to characterize antimicrobial resistance genes across the bacterial communities. qPCR was performed to quantify Mh genome copies, and 16S rRNA gene sequencing was used to assess the broader respiratory microbiome. TE sequencing identified Mh in 100% of TE-tested samples and classified multiple GSVs in all but 3 of 121 samples. GSV profiles clustered within housing groups and varied across cattle populations, indicating structured strain-level diversity. In contrast, Mannheimia spp. were detected in only 47.7% of samples by 16S rRNA sequencing. These findings demonstrate that TE sequencing enables sensitive, strain-level characterization of Mh in cattle and environmental samples and reveals substantial within-population genomic diversity not captured by conventional approaches.IMPORTANCETarget-enriched shotgun sequencing enabled sensitive, strain-level detection of Mannheimia haemolytica (Mh), revealing multiple co-circulating genomic sequence variants (GSVs) within and among cattle groups. This demonstrates greater genetic variability of Mh populations in beef cattle than has been previously recognized. The clustering of GSVs within housing groups, together with the overlap between respiratory and environmental samples, is consistent with the hypothesis that contagious transmission contributes to Mh ecology. These results highlight the potential utility of composite nasal swab and environmental samples for future studies evaluating relationships between Mh genomic variation and disease risk.

Animals

Identification and full genome sequencing of previously unknown sandfly-borne phleboviruses using a newly established capture-based next-generation sequencing approach.

Sandfly-borne phleboviruses cause febrile illness and neuroinvasive disease in humans. While infections are reported in the Mediterranean region, the discovery of previously unknown phleboviruses in sandflies from Kenya suggests a wider geographic distribution. Detection and characterization of novel phleboviruses are often hindered by low-quality and low-viral-load samples. We developed a capture-based target enrichment next-generation sequencing approach that showed a 99%-100% fold enrichment of viral genomes from primary material and provides a robust tool for generating complete genomes of both known and previously unknown viruses. From a collection of 15,652 sandflies in Kenya, we recovered seven complete coding sequences of Embossos, Bogoria, and Kiborgoch viruses, and of two previously unknown phleboviruses, which were named Sosoik and Shable viruses. Sosoik virus shared 83% amino acid identity in its RdRp gene with that of Bogoria virus, while Shable virus shared ca. 88% amino acid identity with viruses of the Salehabad serocomplex. Additionally, a reassortant of Shable virus was detected that possessed an M segment from an undescribed Ponticelli-like virus. DNA barcoding of blood-fed sandflies revealed several potentially novel Sergentomyia species and evidence of host-feeding on humans, livestock, and reptiles, suggesting possibilities for zoonotic transmission. Overall, our findings increase the known genetic diversity of Old World sandfly-borne phlebovirus species from 18 to 25 (by 38.9%), including the detection of viruses from all pathogenic sandfly-borne phlebovirus serocomplexes in East Africa, opening new horizons in disease ecology research.IMPORTANCEKnowledge of the genetic diversity of circulating pathogens is crucial for providing appropriate diagnostics and disease management. This study established a novel capture-based target enrichment next-generation sequencing approach that enabled the near-complete viral genome recovery from primary samples, while native NGS yielded negative or poor-quality results. In addition to the five recently discovered sandfly-borne phleboviruses in Kenya, two previously unknown phleboviruses were detected in sandflies from the same region. The viruses were detected in several sandfly species, which showed diverse host-feeding behaviors, including mixed feeding on humans and chickens. The study significantly advances the understanding of sandfly-borne phleboviruses by uncovering their broader geographic distribution and genetic diversity, particularly in East Africa, highlighting the importance of expanding surveillance efforts beyond traditionally studied regions.

Phlebovirus

Comparative evaluation of probe-capture and conventional metagenomic sequencing across multiple clinical sample types, with analysis of paired bronchoalveolar lavage fluid and blood samples.

Conventional metagenomic next-generation sequencing (mNGS) suffers from host nucleic acid interference and poor performance in low-biomass samples. Probe-capture metagenomic sequencing (PC-mNGS), which enriches microbial targets via hybridization probes, shows superior sensitivity but lacks systematic multi-sample evaluations. This study compared PC-mNGS and mNGS across diverse clinical specimens (bronchoalveolar lavage fluid [BALF], blood, cerebrospinal fluid [CSF]) and assessed the clinical utility of pathogen co-detection in paired BALF-blood samples from sepsis patients. A total of 282 samples (81 BALF, 141 blood, 25 CSF, 35 others) sequenced by both PC-mNGS and mNGS were analyzed. Additionally, 621 paired BALF-blood samples from sepsis patients with pulmonary infections were evaluated. PC-mNGS achieved higher pathogen detection rates (66.67% vs 57.10%, P = 0.000198) than mNGS, particularly in blood (66.67% vs 47.52%, P = 2.5 × 10⁻⁵). PC-mNGS detected more bacteria (19 species exclusive) and fungi (11 species exclusive) than mNGS. Viruses showed comparable detection. BALF and CSF exhibited high overall agreement (OPA: 96.30% and 88%, respectively), while blood had lower concordance (NPA: 54.05%, OPA: 70.92%). A total of 60.55% of BALF-positive samples (PC-mNGS) had co-detected pathogens in blood. Gram-negative bacteria (e.g., Klebsiella pneumoniae) and fungi (e.g., Candida albicans) showed higher blood co-detection rates than viruses. In this study, PC-mNGS detected more pathogens and showed a higher positivity rate than mNGS in blood samples. BALF sequencing data, particularly bacterial reads per million (RPM), may predict bloodstream co-detection, aiding in sepsis management. However, clinical validation and integration with traditional diagnostics are needed to confirm utility. This study highlights PC-mNGS as a promising tool for complex infections but underscores the need for rigorous multi-context validation.IMPORTANCEAccurate and rapid identification of pathogens is critical for effective treatment of severe infectious diseases, such as sepsis. This study demonstrates that probe-capture metagenomic sequencing (PC-mNGS) detected more pathogens in blood samples compared to conventional metagenomic sequencing, especially for bacterial and fungal infections. By analyzing paired lung and blood samples, we show that high pathogen levels in lung fluid may predict bloodstream infection, offering a potential early warning for clinicians. These findings support the use of PC-mNGS as a more sensitive diagnostic tool, which could lead to faster, more targeted therapies and better outcomes for patients with complex infections.

Humans

Next-Generation Sequencing Methods for Sensitive Hepatitis B Viral Genome Analysis: A European Study.

This multicentre study investigated the utility of next-generation sequencing (NGS) to detect and generate hepatitis B virus (HBV) genomes in samples of low viral load (from 0.2 to 6207 IU/mL). 23 HBV DNA-positive plasma samples of genotypes A-E and one HBV-negative control sample were assayed blindly via 9 established NGS methods from 6 European laboratories. Methods included untargeted metagenomics, pre-enrichment by probe-capture followed by Illumina sequencing, and HBV-specific PCR pre-amplification followed by sequencing with Nanopore or Illumina. Full HBV genomes were obtained only from samples with viral loads > 1000 IU/mL using probe-capture methods, > 200 IU/mL using PCR-Illumina methods, > 10 IU/mL using PCR-Nanopore methods, and in no samples using metagenomic methods. Contamination was observed in the negative control and samples with very low viral loads in PCR-based methods. Probe-capture and metagenomic methods detected additional viruses not routinely screened in blood donations, including polyomaviruses and herpesviruses; positive results were confirmed by PCR. In conclusion, NGS may delineate whole-genome sequences at low viral loads if supported by a PCR pre-amplification step. Probe-capture methods also reliably detect HBV without pre-amplification but show limited genome coverage for samples with low viral loads; they may additionally detect a wide range of blood-borne viruses.

Humans

A probe-based capture enrichment method for detection of A-to-I editing in low abundance transcripts.

Exactly two decades ago, the ability to use high-throughput RNA sequencing technology to identify sites of editing by ADARs was employed for the first time. Since that time, RNA sequencing has become a standard tool for researchers studying RNA biology and led to the discovery of RNA editing sites present in a multitude of organisms, across tissue types, and in disease. However, transcriptome-wide sequencing is not without limitations. Most notably, RNA sequencing depth of a given transcript is correlated with expression, and sequencing depth impacts the ability to robustly detect RNA editing events. This chapter focuses on a method for enrichment of low-abundance transcripts that can facilitate more efficient sequencing and detection of RNA editing events. An important note is that while we describe aspects of the protocol important for capturing intron-containing transcripts, this probe-based enrichment method could be easily modified to assess editing within any low-abundance transcript. We also provide some perspectives on the current limitations as well as important future directions for expanding this technology to gain more insights into how RNA editing can impact transcript diversity.

RNA Editing

Innovations in Transgene Integration Analysis: A Comprehensive Review of Enrichment and Sequencing Strategies in Biotechnology.

Understanding the integration of transgene DNA (T-DNA) in transgenic crops, animals, and clinical applications is paramount for ensuring the stability and expression of inserted genes, which directly influence desired traits and therapeutic outcomes. Analyzing T-DNA integration patterns is essential for identifying potential unintended effects and evaluating the safety and environmental implications of genetically modified organisms (GMOs). This knowledge is crucial for regulatory compliance and fostering public trust in biotechnology by demonstrating transparency in genetic modifications. This review highlights recent advancements in T-DNA integration analysis, specifically focusing on targeted DNA enrichment and sequencing strategies. We examine key technologies, such as polymerase chain reaction (PCR)-based methods, hybridization capture, RNA/DNA-guided endonuclease-mediated enrichment, and high-throughput resequencing, emphasizing their contributions to enhancing precision and efficiency in transgene integration analysis. We discuss the principles, applications, and recent developments in these techniques, underscoring their critical role in advancing biotechnological products. Additionally, we address the existing challenges and future directions in the field, offering a comprehensive overview of how innovative DNA-targeted enrichment and sequencing strategies are reshaping biotechnology and genomics.

Transgenes

EvoSNR-Prom: Predicting promoters at single-nucleotide resolution with label-aware transfer learning of the pretrained EVO model.

The precise identification of promoters is crucial for understanding gene regulation. Deep learning methods have achieved considerable success in promoter prediction, yet most operate at the sequence level with coarse-grained labels. This means they label an entire DNA segment as either a "promoter" or "non-promoter," which results in a lack of the nucleotide-level resolution in prediction. In this study, we propose EvoSNR-Prom, a model designed for promoter prediction at single-nucleotide resolution. EvoSNR-Prom is built on the Evo foundation model and formulates promoter identification as a token-level sequence labeling problem, analogous to named entity recognition in natural language processing. To address the limited contextual information available in single-nucleotide tokenization, we introduce a lexicon-enhanced embedding strategy that incorporates biologically meaningful DNA lexicons, enriching contextual representations and improving the model's ability to capture complex sequence motifs. Furthermore, to enhance predictive performance on small size datasets, we integrate a label-aware transfer learning framework to leverage knowledge from well-annotated source species to a target organism. The results across various prokaryotic datasets show that EvoSNR-Prom achieves excellent performance. This work provides a valuable computational framework for the high-precision analysis of gene regulatory elements, contributing to the advancement of promoter prediction at single-nucleotide resolution.

Promoter Regions, Genetic

Evaluation of bone preparation approaches using length-based analysis and targeted sequencing for forensic human identification of historic skeletal remains.

Advances in DNA technology have significantly enhanced the forensic community's ability to develop genetic profiles from unidentified human skeletal remains. However, sampling requires mechanical grinding of hard tissues before DNA isolation. This processing can compromise genetic profiles, particularly in aged bones. We compared the industry-standard pulverization method with an alternative powder-free preparation involving prolonged demineralization and subsequent slicing of 19th-century cortical bone. Data from DNA quantification, STR genotyping, and targeted SNP sequencing were used to evaluate powdered samples versus demineralized slices from paired human bones. Average human DNA yields for pulverized samples and demineralized slices were 0.032&#x2009;ng and 0.692&#x2009;ng, respectively. Demineralized slices recovered more amplifiable DNA than traditional homogenization methods (p&#x2009;<&#x2009;0.05). No pulverized samples produced STR profiles, whereas demineralized slices from the same bone samples yielded partial profiles. Samples underwent DNA repair, library preparation, and hybridization capture using the FORensic Capture Enrichment (FORCE) panel. Applying low-coverage (1X) analysis of high-throughput sequencing (HTS) data, demineralized slices outperformed those prepared by traditional pulverization methods (p&#x2009;<&#x2009;0.05) and substantially increased the information recovered compared with conventional STR analysis methods. Based on HTS data from pulverized samples, DNA fragment length ranged from 27 to 95&#x2009;bp, and FORCE SNP recovery was 33.23%. In contrast, for demineralized slices, DNA fragment length ranged from 85 to 114&#x2009;bp, and FORCE SNP recovery was 83.24%. The required reagents and equipment are typically available in forensic labs, and the workflow outlined herein significantly increases the success of DNA recovery from challenging skeletal samples.

Humans

Targeted Next-Generation Sequencing in Rare Diseases.

Targeted next-generation sequencing (NGS) in rare disease focuses on genetic analysis of specific regions in genome that are linked to a rare disease. In addition to library preparation, sequencing, and data analysis, targeted NGS includes an additional step of target enrichment of selected genes and regions. It allows for more sensitive and profound sequencing, as it is a fast and cost-effective approach with less data burden and is therefore often a method of choice for identifying rare variants in known genes, especially in diagnostics of rare diseases. Several in silico tools address the pathogenicity predictions of rare variants of unknown significance (VUS) and can therefore facilitate clinical interpretation.

Rare Diseases

Hybridization capture increases on-target nanopore sequencing of plant RNA tobamovirus- derived cDNA libraries.

High-throughput sequencing (HTS) can support plant virus surveillance, but host nucleic acids often reduce on-target read recovery. We evaluated a targeted hybridization-capture workflow in which barcoded double-stranded cDNA (ds-cDNA) libraries generated from plant RNA extracts spiked with lyophilized tobamovirus-positive controls were enriched before Oxford Nanopore sequencing. Biotinylated probes targeted conserved regions of cucumber green mottle mosaic virus (CGMMV), species Tobamovirus viridimaculae; pepper mild mottle virus (PMMoV), species Tobamovirus capsici; and tobacco mosaic virus (TMV), species Tobamovirus tabaci. Across four pairs per virus, relative target-read abundance increased after capture from 0.76 &#xb1; 0.33% to 37.62 &#xb1; 15.72% for CGMMV, 8.16 &#xb1; 3.86% to 24.68 &#xb1; 12.34% for PMMoV, and 15.62 &#xb1; 10.40% to 36.83 &#xb1; 30.33% for TMV. Exact two-sided Wilcoxon signed-rank tests yielded P = 0.125 for each virus; with four nonzero differences in a common direction, this was the minimum attainable two-sided P value. Genome-coverage breadth was maintained. Retrospective duplex qPCR supported an increased virus-to-18S ratio for CGMMV, showed a variable PMMoV response, and showed a decreased virus-to-18S ratio for TMV because the 18S signal shifted earlier by as much as or more than the TMV signal. The findings provide proof-of-concept evidence for target-dependent library enrichment but do not establish analytical sensitivity, diagnostic performance, or field validity. Validation with naturally infected, low-titer, and mixed-infection samples and comparison with simpler targeted workflows are required.

biosecurity

Genome-wide profiling the integration patterns with T7-PCR.

Integration of exogenous gene fragments into the host genomes is a widely used and powerful method for studying gene functions, advancing molecular breeding, and conducting gene therapy. Accurately identifying the integration sites is essential for ensuring both the safety and efficacy of genome engineering efforts. However, current mapping techniques are constrained by high costs and a low signal-to-noise ratio. In this study, we developed an innovative tool for mapping integration sites, leveraging T7 polymerase-mediated in vitro transcription (T7-IVT) to capture the junction fragments surrounding integration loci. This approach converts genomic flanking sequences into RNA, enabling the simultaneous enrichment of junction fragments and the elimination of background genomic DNA, thereby significantly enhancing the signal-to-noise ratio. We have validated the efficiency of this method, named T7-PCR, across yeast, plant, and human cells under diverse integration scenarios. T7-PCR outperforms current next-generation sequencing (NGS)-based mapping strategies in terms of efficiency and accuracy, with minimal positional effects. This method is highly applicable for high-throughput transgene screening and also supports the development of next-generation tools for targeted integration of large fragments.

Humans

Inferring the sensitivity of wastewater metagenomic sequencing for early detection of viruses: a statistical modelling study.

BACKGROUND: Metagenomic sequencing of wastewater (W-MGS) can in principle detect any known or novel pathogen in a population. We aimed to quantify the sensitivity and cost of W-MGS for viral pathogen detection by jointly analysing W-MGS and epidemiological data for a range of human-infecting viruses. METHODS: In this statistical modelling study, we analysed sequencing data from four studies of untargeted W-MGS to estimate the relative abundance of 11 human-infecting viruses. Corresponding prevalence and incidence estimates were obtained or calculated from academic and public health reports. We combined these estimates using a hierarchical Bayesian model to predict relative abundance at set prevalence or incidence values, allowing comparison across studies and viruses. These predictions were then used to estimate the sequencing depth and concomitant cost required for pathogen detection using W-MGS with or without use of a hybridisation capture enrichment panel. FINDINGS: After controlling for variation in local infection rates, relative abundance varied by orders of magnitude across studies for a given virus. For instance, a local SARS-CoV-2 weekly incidence of 1% corresponded to a predicted SARS-CoV-2 relative abundance ranging from 3&#xb7;8&#x2009;&#xd7;&#x2009;10-10 to 2&#xb7;4&#x2009;&#xd7;&#x2009;10-7 across studies, translating to orders-of-magnitude variation in the cost of operating a system able to detect a SARS-CoV-2-like pathogen at a given sensitivity. Use of a respiratory virus enrichment panel in two studies greatly increased predicted relative abundance of SARS-CoV-2, lowering yearly costs by 27-fold (from US$7&#xb7;87 million to $287&#x2009;000) and 29-fold (from $1&#xb7;98 million to $69&#x2009;100) for a system able to detect a SARS-CoV-2-like pathogen before reaching 0&#xb7;01% cumulative incidence. INTERPRETATION: The large variation in viral relative abundance after controlling for epidemiological factors indicates that other sources of inter-study variation, such as differences in sewershed hydrology and laboratory protocols, have a substantial impact on the sensitivity and cost of W-MGS. Well chosen hybridisation capture panels can greatly increase sensitivity and reduce cost for viruses in the panel, but might reduce sensitivity to unknown or unexpected pathogens. FUNDING: The Wellcome Trust, Open Philanthropy, and Musk Foundation.

Humans

Optimizing a culture-enriched hybrid metagenomics pipeline to assess the AMR footprint of livestock manure in anaerobic digestate.

The role of environmental samples from livestock production systems, including manure and anaerobic digestate, as reservoirs of antimicrobial resistance genes (ARGs) is likely underestimated because conventional metagenomic approaches can overlook low-abundance ARGs and often lack the resolution to associate these genes with their microbial hosts and co-localized mobile genetic elements (MGEs). We evaluated whether culture-enriched metagenomics (CEMG), with and without antibiotic selection, enhances ARG detection in anaerobic digestate and improves the resolution of ARG-MGE-host associations using hybrid short- and long-read metagenomic assembly. CEMG increased ARG recovery; mean ARG abundance rose from 15.4 counts per million (CPM) in metagenomic fresh digestate (FD) to 124 CPM in CEMG without antibiotics and 160 CPM in antibiotic-selective CEMG. In FD, only 9 unique ARGs were detected, whereas CEMG recovered 112, including ARGs of clinical importance, such as glycopeptide resistance, beta-lactamase genes, and the cfr 23S rRNA methyltransferase conferring cross-resistance to multiple antibiotic classes. Antibiotic selection induced targeted, class-specific shifts in ARG profiles, with ARGs associated with tetracycline resistance consistently enriched across treatments. Hybrid metagenomic assembly resolved the genomic context of 784 ARGs, of which 59.3% were co-localized with at least one class of MGEs, predominantly plasmids and integrative conjugative elements/integrative mobilizable elements. Biocide and metal resistance genes frequently co-occurred with ARGs on the same contigs. Together, these findings demonstrate that antibiotic-selective culture enrichment enhances resistome surveillance by improving detection of low-abundance ARGs, while hybrid assembly provides critical genomic context for assessing their mobility and host associations.IMPORTANCELivestock manure and its byproducts, such as anaerobic digestate, are recognized as important environmental reservoirs of antimicrobial resistance genes (ARGs) and resistant bacteria, yet current metagenomic approaches may underestimate this risk by failing to detect low-abundance but clinically relevant ARGs. Here, we show that integrating culture enrichment with hybrid metagenomics improves ARG recovery and reveals ARG co-localization with mobile genetic elements and putative bacterial hosts. This approach captures a cultivable and condition-responsive fraction of the resistome that is not readily accessible through direct metagenomic sequencing alone, providing a more informative framework for environmental AMR surveillance.

anaerobic digestion

Target Capture of Ancient Shell DNA Enables Phylogenetic Reconstruction of Deep-Sea Molluscs.

Target capture is widely used to enrich endogenous DNA from calcium phosphate skeletal material in vertebrates, but its performance on calcium carbonate hard parts widely produced by invertebrates remains poorly understood. Here, we compared DNA recovery from four fresh and 12 ancient (eight radiocarbon-dated to 1671-1135&#x2009;years old before present) deep-sea vesicomyid clam shells, including species Archivesica marissinica, A. nanshaensis and A. okutanii, using whole-genome sequencing (WGS) or target capture of ultraconserved elements (UCEs). WGS achieved 16.65% on-target read recovery of UCEs from fresh soft tissue, but <&#x2009;1% from shell specimens. By contrast, UCE capture in the same specimen increased on-target reads by up to 155-fold, reaching 29.84% in fresh shells and up to 72-fold, reaching 19.89% in ancient shells. Target capture of UCEs recovered 142-1001 loci per sample compared to 0-230 with WGS alone. Ancient shells of A. marissinica and A. okutanii, based on reads mapped with bwa-mem2 and bbmap, exhibited characteristic post-mortem DNA damage signals, with average 5'-end C-to-T misincorporation rates of 3.46% and 15.97%, respectively, exceeding the levels observed in fresh A. marissinica shells (maximum 1.24%). UCE-based phylogenetic reconstructions incorporating shell ancient DNA recovered two major clades within Pliocardiinae, consistent with published phylogenomic trees. Together, these findings demonstrate that target-capture enrichment enables effective recovery of highly degraded DNA from ancient mollusc shells and supports robust phylogenetic inference at the intrageneric scale, expanding the utility of shells-one of the most abundant invertebrate remains-for evolutionary, biogeographic and conservation studies.

Animals

Molecular DNA enrichment methods for parasite genomic sequencing in clinical samples: a systematic review.

Parasitic diseases such as malaria, Chagas disease, leishmaniases, and helminthiases are major causes of sickness and death in low- and middle-income countries. The high genetic diversity of these pathogens affects virulence, immune evasion, and diagnostic accuracy. Although Whole Genome Sequencing (WGS) is a powerful tool for tracking genetic variants and drug resistance, low parasitemia and the predominance of host DNA limit its application to clinical samples. This study systematically reviewed molecular strategies to improve the recovery of parasite DNA from clinical samples, following PRISMA 2020 guidelines and registered in PROSPERO. Searches of PubMed, Scopus, Web of Science, and LILACS up to December 2025 identified 20 eligible studies, most of which focused on protozoa, particularly Plasmodium spp. The main approaches included hybridization capture, selective whole-genome amplification, host DNA depletion, and in silico enrichment via adaptive sampling. Overall, no single method is suitable for all parasites analyzed; the optimal approach depends on the pathogen, sample type, and research objective. The review emphasizes that parasite DNA enrichment is essential for enabling WGS in clinical settings, underscoring the need for protocol standardization and cost-effectiveness analyses to support public health genomic surveillance.

Adaptive sampling

Comparison of paralog identification methods and their impact on species tree topologies in target capture phylogenomics within the Sindora clade (Detarioideae: Leguminosae).

Target capture is a common method of generating high throughput DNA sequencing data for phylogenetic reconstruction of species relationships, for which single copy genes are usually most informative. However, a pervasive problem with target capture is that putatively single copy genes may in fact be paralogs resulting from gene duplication, which are problematic for phylogenetic inference because their evolutionary history may differ from the divergence history of species. Here, we use as a case study a target enrichment dataset of 88 species of Detarioideae (Leguminosae) with a focus on the Sindora clade to examine approaches for handling paralogs, including the built-in paralog handling functions in HybPiper and CAPTUS, plus subsequent steps using Putative Paralog Detection and the tree-based Yang & Smith orthology inference approach. We compare the paralogs flagged using these methods and verify their performance with BLAST mapping against a reference genome sequence of Sindora glabra, and then subsequently compare the species tree topologies produced across these methods. Our comparisons of paralogs flagged across the Sindora clade show that the Putative Paralog Detection pipeline was the most accurate in identifying paralogs in terms of its similarity to the BLAST mapping, followed by the built-in paralog identification function of CAPTUS. However, the results we recovered for the Detarioideae subfamily suggest that the largest differences in species tree topology resulted from the use of paralog-filtered alignments (such as with the Putative Paralog Detection pipeline and the Yang & Smith orthology inference approaches) rather than just by removing the sequences of identified paralogous genes. This was the true for HybPiper-assembled datasets but was not seen in CAPTUS-assembled datasets. In all comparisons, the topological differences caused by different paralog handling methods tended to be confined to clades where processes such as hybridisation and introgression are prevalent. Our study provides a roadmap to establish the best approach to identify, eliminate or separate paralogs in the absence of a chromosomally contiguous reference genome for a study group, and highlights the importance of careful data inspection and processing in addition to understanding the extent of paralogy and paralog characteristics (e.g. sequence divergence between copies) for their study group.

Phylogeny