Search PubMedSearch

SEARCH · Search PubMed

Results for “Sequencing quality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

SeqQC-former: A sequence-quality fusion framework for QC-aware review prioritization of candidate somatic SNVs in cancer genomics.

The accurate prioritization of candidate somatic single-nucleotide variants (SNVs) remains a challenge due to the substantial variability in sequencing quality across genomic loci. SeqQC-Former is a sequence-quality fusion framework that integrates the local nucleotide context with read-level quality-control (QC) covariates derived from matched tumor-normal sequencing data. This integration generates QC-aware prioritization scores for the downstream review of candidate variants. Unlike conventional variant callers, SeqQC-Former is designed not to infer biological truth but to support post-calling review and prioritization under heterogeneous sequencing conditions. The framework was trained and evaluated on a SEQC2-derived dataset comprising 89,447 candidate loci, including 1378 positive and 88,069 negative loci. In chromosome-held-out validation, which aims to reduce potential genomic-position leakage, SeqQC-Former demonstrated strong discrimination (AUROC = 0.9479; AUPRC = 0.9448), indicating good generalization to previously unseen chromosomes. Given that the SEQC2-derived labels contain QC-associated information; these results should be interpreted as an evaluation of QC-aware prioritization capability rather than an independent validation of biological variant correctness. Ablation analyses revealed that structured QC covariates provided the dominant predictive signal under the current SEQC2-derived labeling regime. SeqQC-Former achieved a significantly higher AUROC than classical machine-learning baselines, as determined by DeLong's test (p&#x202f;<&#x202f;0.01). Application to 53,164 glioblastoma variants demonstrated that external predictions were sensitive to QC scaling and threshold selection, underscoring that model outputs should be interpreted as QC-dependent prioritization scores rather than calibrated probabilities or definitive biological classifications. Overall, SeqQC-Former offers a reproducible post-calling QC-aware prioritization framework for large-scale somatic SNV review and underscores the importance of explicitly modeling sequencing-quality information when interpreting structured cancer genomics datasets.

Humans

The accuracy of DNA sequences: estimating sequence quality.

In this paper we describe a method for the statistical reconstruction of a large DNA sequence from a set of sequenced fragments. We assume that the fragments have been assembled and address the problem of determining the degree to which the reconstructed sequence is free from errors, i.e., its accuracy. A consensus distribution is derived from the assembled fragment configuration based upon the rates of sequencing errors in the individual fragments. The consensus distribution can be used to find a minimally redundant consensus sequence that meets a prespecified confidence level, either base by base or across any region of the sequence. A likelihood-based procedure for the estimation of the sequencing error rates, which utilizes an iterative EM algorithm, is described. Prior knowledge of the error rates is easily incorporated into the estimation procedure. The methods are applied to a set of assembled sequence fragments from the human G6PD locus. We close the paper with a brief discussion of the relevance and practical implications of this work.

Algorithms

TargetQC: A targeted quality control framework for clinical genomic testing.

Reliable genetic testing depends on accurate assessment of sequencing quality in clinically relevant genomic regions that directly influence variant interpretation. We developed TargetQC, a flexible quality control framework that supports user-defined gene sets, coverage thresholds, and variant sets for evaluating sequencing performance across exome sequencing (ES) and genome sequencing (GS) platforms. TargetQC assesses exon and gene coverage, identifies regions meeting predefined coverage thresholds, evaluates variant detection accuracy, and measures sequencing quality at pathogenic variant sites. We applied TargetQC to the reference sample NA12878 and 665 clinical samples across five ES platforms and one GS platform. ES-VendorB and ES-VendorE achieved the most complete coverage of OMIM coding regions in NA12878, whereas ES-VendorD and ES-VendorE showed the highest coverage compliance in clinical samples. ES-VendorB and GS demonstrated the highest variant detection accuracy. TargetQC provides a practical framework for benchmarking sequencing performance and informing platform selection in clinical genomics.

exome sequencing

A hybrid and cost-efficient barcoding strategy for full-length 16S rRNA gene nanopore sequencing of environmental samples.

BACKGROUND: Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford Nanopore Technology (ONT) long-read sequencing to full-length 16S rRNA amplicon in samples from natural soil amended with lignocellulosic biomass and a simplified microbial community derived from cultures grown on selective and differential carboxymethyl cellulose (CMC)-based substrates, with the aim to evaluate the difference in performance between a real, complex community and a less complex system. To reduce consumable costs, we substituted the standard ONT Barcoding kits with an in-house hybrid barcoding workflow. Specifically, PacBio PCR-based barcoding protocol was used for sample indexing, followed by library preparation using the ONT Ligation Sequencing Kit. This simplified approach retained compatibility with MinION and Flongle flow cells and supported accurate downstream demultiplexing while lowering barcode costs substantially. Additionally, a new bioinformatic workflow tailored to ONT data was implemented. RESULTS: Overall, the hybrid protocol significantly reduced per-sample barcoding costs while preserving high sequencing quality and throughput. The sequencing run yielded over 5 Gb of quality-filtered data (Q-score &#x2265; 10). Furthermore, the new bioinformatic workflow allowed taxonomic assignment at the species level for 49.38% of annotated taxa, compared to just 4.59% using Illumina NovaSeq sequencing of the V3-V4 region. ONT also recovered 2.3 times more genera and 1.3 times more families. Although 16S rRNA gene sequencing often cannot distinguish between closely related species, particularly within taxonomically complex groups, in this work, full-length reads substantially improved both taxonomic resolution and database matching. CONCLUSIONS: These results show that full-length 16S rRNA sequencing with ONT, paired with a low-cost barcoding strategy, enhanced taxonomic resolution compared to short-read workflows. This approach also offers a scalable and cost-effective option for high-resolution microbiome profiling in research and applied settings.

RNA, Ribosomal, 16S

Global Profiling and Analysis of 5' Monophosphorylated mRNA Decay Intermediates.

During RNA turnover, the action of endo- and exo-ribonucleases can yield RNA decay intermediates with specific 5' ends. These RNA decay intermediates have been demonstrated to be the outcome of decapping, microRNA-directed endo-cleavage, or the protected fragments of ribosomes and exon-junction complexes. Therefore, global analysis of RNA decay intermediates can facilitate studies of many RNA decay pathways. In this chapter, we describe a high-throughput sequencing protocol named parallel analysis of RNA ends (PARE), which allows genome-wide profiling of 5' monophosphorylated mRNA decay intermediates from plants or other eukaryotes. Also, we present the tools and scripts necessary for the proper analysis of RNA degradome data obtained from the PARE method. Details and modifications of library construction procedures and bioinformatic analyses to optimize sequencing quality and cope with emerging sequencing platforms and findings are highlighted.

RNA Stability

Molecular Epidemiology of Non-Polio Enterovirus: Insights From L20B Cell Line Adaptation From Children With Acute Flaccid Paralysis in Pakistan.

BACKGROUND: Non-polio enteroviruses (NPEVs) are increasingly implicated in acute flaccid paralysis (AFP), often resembling poliomyelitis and complicating eradication efforts. In Pakistan, limited molecular surveillance has hindered comprehensive characterization. The L20B cell line, designed for poliovirus detection, occasionally supports NPEV replication, challenging AFP case interpretation. METHODS: Between January 2021 and December 2022, 4615 stool samples from AFP cases in children &#x2264;15 years were analyzed. Of these, 435 were identified as NPEVs via L20B cytopathic effects and intertypic differentiation reverse transcription-polymerase chain reaction. VP1 sequencing was performed on 218 representative isolates, yielding 153 high-quality sequences (70.2%). The 224/222 primer set showed superior amplification. Phylogenetic analysis used MUSCLE alignment and the Neighbor-Joining method in MEGA X, with statistical evaluation of epidemiological data. RESULTS: NPEVs were frequently found in L20B-positive AFP cases, highlighting the cell line's limited specificity. Most cases involved children under 5, with a slight male bias. Enterovirus B was predominant (98.0%), especially Echovirus 7 (20.3%) and Echovirus 11 (10.5%), followed by Coxsackievirus B1 and Echovirus 33 (5.9% each). Geographic clustering was noted in Punjab (45.1%), Khyber Pakhtunkhwa (30.7%) and Sindh (20.3%), with seasonal peaks in late summer and early autumn. Phylogenetic data revealed localized Enterovirus B circulation with minimal genetic variation. CONCLUSIONS: The detection of diverse NPEVs in L20B-positive AFP cases emphasizes their relevance in post-polio surveillance. Incorporating routine VP1 sequencing, optimized primer use, and targeted seasonal and regional monitoring is vital to reduce diagnostic uncertainty and inform public health strategies.

Humans

Assessing Hardy-Weinberg equilibrium in T2T-aligned 1000 genomes project.

Quality control of markers in genome-wide association studies often includes testing for Hardy-Weinberg equilibrium (HWE). However, this is usually implemented in a homogeneous population without stratifying by sex. Previous work indicates sex-based selection at numerous autosomal loci in cohorts with active recruitment. Sex chromosome sequences can also interfere with autosomal SNPs. These motivate a re-examination of HWE in sex-aware analyses. Using the telomere-to-telomere (T2Tv2)-aligned high-coverage whole genome sequencing data from 2,490 individuals in the 1000 Genomes Project, we examined genome-wide sex-specific deviations from HWE across five super-populations. Our analyses were restricted to bi-allelic SNPs with non-missing genotypes and minor allele frequency (MAF) &#x2265;5% in both sexes of the five super-populations. We applied an allele-based framework to quantify both the magnitude and direction of Hardy-Weinberg disequilibrium (HWD), followed by a second-order omnibus meta-analysis that combined HWD results across populations and sexes. At a genome-wide significance threshold of p&#x2009;<&#x2009;5e-8, 0.9% of autosomal SNPs exhibited significant deviations from HWE. The majority of these deviations were associated with genomic features indicative of poor sequence quality. Restricting the analysis to reliable genomic regions substantially reduced the number of signals, yielding 255 autosomal SNPs and one non-pseudoautosomal chromosome X SNP. Among these, 140 autosomal SNPs displayed significant heterogeneity across populations but not across sexes. Notably, eight SNPs within a 15-bp region on chromosome 14q31.3 showed excess heterozygosity in both sexes of the African super-population (AFR). Finally, we developed a multivariate predictor of HWD based on sequence features, providing a practical tool that can be integrated into existing quality control pipelines for whole genome sequencing studies.

Journal Article

Real-world emergence of nirsevimab resistance in breakthrough infections with respiratory syncytial virus-B: a multicentre observational study in France.

BACKGROUND: Respiratory syncytial virus (RSV) is a leading cause of lower respiratory tract infection in infants. Nirsevimab, a long-acting monoclonal antibody targeting a conserved epitope on the prefusion F protein (site &#x3a6;), has shown high efficacy in clinical trials and early real-world studies. Although widespread resistance has not been reported, concerns remain about the emergence of escape variants, particularly among RSV-B viruses. During the 2024-25 RSV season in France, RSV-B predominated, providing a unique opportunity to examine breakthrough infections with RSV-B and resistance at a large scale. The study aimed to characterise RSV escape from nirsevimab using genotypic and phenotypic methods. METHODS: This POLYRES-2 project was a multicentre, national, observational study conducted in hospital settings (inpatients and outpatients) across France during the 2024-25 RSV season. We included infants aged 1 year or under with a RT-PCR-confirmed RSV infection in routine care, regardless of whether they had received nirsevimab. Infants were identified through hospital virology laboratory databases. Each participating centre was requested to include a balanced number of nirsevimab-exposed and non-exposed infected infants throughout the study period. Clinical data were retrieved from electronic medical records. We compared RSV susceptibility to nirsevimab in infants who received nirsevimab with that in nirsevimab-naive infants. Respiratory samples were sequenced for full-length RSV genomes. To ensure reliability, phylogenetic and mutational analyses were restricted to high-quality sequences with greater than or equal to 90% genome coverage and complete reads across the nirsevimab-binding site. Clinical RSV isolates were tested for neutralisation by nirsevimab. We analysed F candidate substitutions using a fusion inhibition assay. The primary outcomes were presence of resistance-associated substitutions (RASs) in the RSV F protein (site &#x3a6;) and phenotypic resistance to nirsevimab. FINDINGS: Among 1023 RSV-infected infants, 858 (83&#xb7;9%) had full-length RSV genome sequences: 419 (48&#xb7;8%) from nirsevimab-treated breakthrough infections (212 [50&#xb7;6%] RSV-A, 207 [49&#xb7;4%] RSV-B) and 439 (51&#xb7;2%) from nirsevimab-naive infants (192 [43&#xb7;7%] RSV-A, 247 [56&#xb7;3%] RSV-B). RASs were identified in two of 195 RSV-A breakthrough infections (1&#xb7;0%) and in 23 of 184 RSV-B breakthrough infections (12&#xb7;5%). In RSV-A, the only RAS was F:K209E, conferring intermediate resistance. In RSV-B, resistance was more frequent and diverse than in RSV-A: 12 of 23 (52.2%) resistant viruses carried a substitution at residue 208 (F:N208D, F:N208I, F:N208K, F:N208S, or F:N208Y). Additional novel substitutions, including F:I64V/F:K65E, F:K68I, F:L204S, and F:P205S, also mediated resistance. Notably, a resistant RSV-B variant (F:N208S) was detected almost 1 year after prophylaxis. No resistant RSV was detected in nirsevimab-naive infants. INTERPRETATION: Resistance to nirsevimab in RSV-B can emerge in real-world settings, affecting around 12% of breakthrough infections and showing greater diversity than previously recognised, although the clinical impact remains constrained by available evidence. Detection of resistant variants long after prophylaxis highlights the need for extended genomic surveillance. Integration of clinical and virological data will be essential to sustain the long-term effectiveness of RSV monoclonal antibody programmes. FUNDING: This study was supported by a grant from the Agence Nationale de Recherche sur le Sida et les h&#xe9;patites virales - Maladies Infectieuses Emergentes and the French Ministry of Health and Prevention.

Humans

Large-scale simulation of coverage and error rate tradeoffs for cancer detection in cell-free DNA whole-genome sequencing.

MOTIVATION: Cell-free DNA (cfDNA) whole-genome sequencing (WGS) is a promising approach for detecting cancer recurrence. It enables cancer detection by identifying all tumor-derived cfDNA (ctDNA) molecules carrying somatic single nucleotide variants (sSNVs). While ideally, a sequencing platform should be highly accurate for reliable ctDNA detection, in reality, all sequencing platforms introduce sequencing errors that generate false positives indistinguishable from true SNVs. Understanding how sequencing parameters influence ctDNA detection sensitivity at low tumor fractions (TFs) in cfDNA samples is essential for guiding sequencing strategies in clinical contexts. To model cfDNA sequencing for tumor detection, which contains asymmetric noise and multiple interacting parameters, analytical modeling is intractable, motivating large-scale parallelized simulation. RESULTS: We developed a simulation framework to generate in silico cfDNA data across 10 cancer types. In total, 480 million cfDNA samples were simulated from tumor WGS profiles. Overall, the lowest detectable TF differs substantially between cancer types under identical sequencing conditions due to variations in mutational load. For cancers with high mutational load, 3&#xd7; coverage with low-error techniques reliably detects TFs below 0.1%. In contrast, cancers with low mutational load require at least six-fold higher coverage to achieve comparable detection thresholds. Increasing sequencing quality scores from Q30 to Q55 at 30&#xd7; coverage further enhances sensitivity, enabling detection of TFs as low as 1&#x2009;&#xd7;&#x2009;10-5. This study provides a comprehensive framework for optimizing sequencing parameters, offering valuable guidance for tailoring future technology development for specific cancer types and clinical applications. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/UMCUGenetics/cfdetect/tree/main.

Whole Genome Sequencing

The Aggregated Gut Viral Catalogue (AVrC): A unified resource for exploring the viral diversity of the human gut.

The growing interest in the role of the gut virome in human health and disease, has led to several recent large-scale viral catalogue projects mining human gut metagenomes each using varied computational tools and quality control criteria. Importantly, there has been to date no consistent comparison of these catalogues' quality, diversity, and overlap. In this project, we therefore systematically surveyed nine previously published human gut viral catalogues. While these catalogues collectively screened >40,000 human fecal metagenomes, 82% of the recovered 345,613 viral sequences were unique to one catalogue, highlighting limited redundancy between the ressources and suggesting the need for an aggregated resource bringing these viral sequences together. We further expanded these viral catalogues by mining 7,867 infant gut metagenomes from 12 large-scale infant studies collected in 9 different countries. From these datasets, we constructed the Aggregated Gut Viral Catalogue (AVrC), a unified modular resource containing 1,018,941 dereplicated viral sequences (449,859 species-level vOTUs). Using computational inference tools, annotations were obtained for each vOTU representative sequence quality, viral taxonomy, predicted viral lifestyle, and putative host. This project aims to facilitate the reuse of previously published viral catalogues by the research community and follows a modular framework to enable future expansions as novel data becomes available.

Humans

ATAC-seq in Emerging Model Organisms: Challenges and Strategies.

The Assay for Transposase-Accessible Chromatin with sequencing (ATAC-seq) is a versatile and widely utilized method for identifying potential regulatory regions, such as promoters and enhancers, within a genome. ATAC-seq has been successfully applied to a wide range of established and emerging model organisms. However, implementing this method in emerging model systems, such as arthropods, can be challenging due to several factors that influence data quality. These factors include the availability of a sufficient amount and quality of tissue or cells, the need for species- and tissue-specific protocol optimization, the completeness and accuracy of the reference genome, and the quality of the genome annotation. In this article, we emphasize the key steps in the ATAC-seq protocol that, based on our experience, have the greatest impact on data quality when adapting this method for emerging model organisms. Specifically, we discuss the importance of nuclei isolation, the incubation conditions of the Tn5 transposase, and PCR amplification of the library. Furthermore, we outline essential quality checkpoints during the bioinformatic analysis of ATAC-seq data to assist in assessing data integrity and consistency. Given that many emerging model organisms may not be readily available in laboratory cultures, we also emphasize the importance of evaluating how different preservation methods affect ATAC-seq data quality. Based on examples in one spider and one ant species, we demonstrate that replication and thorough quality controls at all steps of the protocol and data analysis are essential to assess the usability of ATAC-seq data. Our data highlights the importance of isolating the right number of intact nuclei, as well as ensuring optimal amplification conditions during library preparation to obtain good-quality sequence data for downstream analyses. We recommend using fresh tissue samples if possible because we show that direct cryopreservation of the tissue may affect chromatin integrity. This effect could be avoided or reduced by preserving the homogenate in cell culture medium. Overall, we explain the ATAC-seq protocol and downstream analyses in detail and give step-by-step advice to researchers who are new to the field and want to implement this method. With careful planning and validation, ATAC-seq can reveal the regulatory landscape of a genome and aid in identifying elements that govern gene expression.

Animals

Transition of Staphylococcus aureus tetracycline resistance plasmid pT181 from independent multicopy replicon to predominantly integrated chromosomal element over 65 years.

Mobile genetic elements (MGEs), including plasmids, phages and genome islands, are major sources of bacterial genetic diversity. The small plasmid pT181 confers tetracycline resistance in bacterial pathogen Staphylococcus aureus via an efflux pump, TetK. pT181 was one of the earliest sequenced S. aureus plasmids, and has been isolated in both clinical and livestock-associated strains for decades, both as an independent replicon and integrated in the chromosome as part of staphylococcal cassette chromosome mec (SCCmec). Bacterial genome analysis tools and high-quality sequences with metadata are publicly available, but these resources remain underleveraged for examining historical data, especially when studying the spread of MGEs across a species and over time. Using publicly available reads and metadata, we explored the evolution of pT181 over almost seven decades of samples to identify temporal trends in sequence evolution, copy number changes, and spread across S. aureus and beyond. pT181 was prevalent across S. aureus (found in 9.5% of 83,366 genomes tested), with a conserved sequence outside of three hypervariable regions. The history of pT181 since 1954 is characterized by spread across strains, significant variation in plasmid copy number of the independent replicon, and increasing frequency of integration of the plasmid into the S. aureus chromosome. We have identified multiple chromosomal integration locations of the plasmid, including outside of the previously characterized SCCmec. We find that pT181 has been transferred across staphylococcaceae and into a Gram-negative species. The repeated integration of pT181 into the chromosome may indicate co-evolution of the plasmid and the host, potentially to facilitate increased antibiotic resistance.

Journal Article

ExoMeth sequencing of DNA: eliminating the need for subcloning and oligonucleotide primers.

A method is reported for sequencing DNA based on exonuclease III digestion and strand protection by using modified nucleoside triphosphates. Up to 10 kilobases of sequence information may be obtained from each strand of a given template without subcloning. Prior knowledge of the restriction map is not important; prior knowledge of any of the sequence is not required. Nor are oligonucleotide primers needed. Double-stranded cosmids, plasmids, lambda phage, or linear molecules (including amplified molecules) may be used as starting material. The method creates a single-stranded template from these starting molecules, thus generating high-quality sequence ladders. Most commonly used DNA polymerases may be utilized, including reverse transcriptase and T7 DNA polymerase. The approach is "ordered", so little time is wasted on redundant sequencing.

Base Sequence

Mobile genetic elements-driven partitions of mega-plasmids resistome in Salmonella Infantis.

Salmonella enterica serovar Infantis (S. Infantis) becomes the primary pathogen among the top Salmonella serotypes, contributing to numerous cases of foodborne illness annually in the United States. S. Infantis infection has spread rapidly worldwide, especially the clones with pESI-like plasmids. However, the underlying mechanisms regarding the transmission of S. Infantis, particularly mobile genetic elements (MGEs), mediated horizontal gene transfer, are limited. The objective of this study was to evaluate the relationship, if any, among MGEs, antibiotic-resistant genes (ARGs), and virulence factors (VFs) within S. Infantis via genomic analysis. A total of 91 S. Infantis complete genomes with high sequencing quality were selected for downstream bioinformatic analysis. The results showed that the majority of VFs were located in the bacterial chromosomes, while most ARGs were carried by S. Infantis mega-plasmids in an MGE-favored manner. Integrons and transposons were closely associated with certain ARGs, but prophages within mega-plasmids displayed a diverse ARG profile. Collectively, MGE-mediated horizontal gene transfer might lead to ARG acquisition by mega-plasmids, subsequently contributing to the resistome of S. Infantis. Our findings provide insights into the development of MGE-associated resistome in S. Infantis that could inform more effective prevention and intervention strategies to control this pathogen, further ensuring public health and safety.IMPORTANCEThe rapid emergence and transmission of antibiotic-resistant foodborne pathogens pose a significant risk to public health, necessitating the discovery of underlying mechanisms to control multidrug-resistant pathogens. Salmonella enterica serovar Infantis (S. Infantis) has become a pathogen of clinical and epidemiological relevance in recent years, ranking as the top prevalent serovar associated with foodborne illnesses and exhibiting resistance to several antibiotics. The current investigation of multidrug resistance (MDR) S. Infantis strains primarily emphasized the presence of mega-plasmids. However, the question of how mega-plasmids contribute to the transmission of antibiotic-resistant genes (ARG) is unaddressed. Utilizing the genomic characterization of S. Infantis complete genomes with high quality, our study revealed that the resistome of S. Infantis mega-plasmids-the primary ARG reservoirs of S. Infantis-followed a specific pattern of mobile genetic elements (MGEs). Monitoring the spread of MGE-carried ARGs within mega-plasmids should be considered in future surveillance.

Interspersed Repetitive Sequences

Phylogenetic inconsistency of pairwise SNP clustering for inferring tuberculosis transmission in a high-burden, endemic setting: a case study from Thailand.

Whole-genome sequence analysis is now widely used to delineate tuberculosis transmission clusters. A standard practice is to cluster bacterial isolates based on a fixed maximum genome-wide pairwise single nucleotide polymorphism (pwSNP) distance threshold. In this study, we evaluated the phylogenetic consistency of pwSNP-distance clustering with thresholds ranging between 1 and 25 single nucleotide polymorphisms (SNPs) using two contrasting data sets: (i) a data set from the UK (N = 390) published by T. M. Walker, C. L. C. Ip, R. H. Harrell, J. T. Evans, et al. (Lancet Infect Dis 13:137-146, 2013, https://doi.org/10.1016/S1473-3099(12)70277-3), which was foundational to the establishment of this method, and (ii) a data set from Thailand (N = 3,341), characterized by persistent transmission and sparse, non-systematic sampling. For the UK data set, the standard pwSNP-distance clustering using thresholds of &#x2265;12 SNPs yielded entirely monophyletic clusters and showed high concordance with a comparative monophyly constrained, tree-based method. In contrast, for the Thai data set, pwSNP-distance clustering often generated non-monophyletic clusters, even by the 25-SNP threshold. The pwSNP-distance and comparative tree-based clustering methods only showed large consistency at thresholds of &#x2265;22 SNPs. This suggests that SNP clusters defined by low distance thresholds (i.e., <12 SNPs for the UK data set, and <22 SNPs for the Thai data set) may lack robustness, and the problem is particularly severe for data sets characterized by persistent transmission, likely due to poorer cluster separation. Moreover, our findings indicate that large cluster sizes, high maximum intra-cluster genetic distances, and broad sample collection time spans may serve as useful indicators of potentially non-monophyletic clusters. We also demonstrate that mixed infections can produce spurious, phylogenetically long-range SNP linkages, underscoring the necessity of strict sequence quality control.IMPORTANCEFixed-threshold pairwise single nucleotide polymorphism (pwSNP)-distance clustering is commonly used to delineate tuberculosis transmission clusters. From an epidemiological perspective, a genuine transmission cluster must be monophyletic, originating from a single source. However, pwSNP-distance clustering is inherently simplistic and can therefore violate this principle, making the assessment of its phylogenetic consistency critical. Our results demonstrate that while this method effectively delineated complete transmission clusters for the data set from the UK, a low-burden and non-persistent transmission setting, it frequently generated non-monophyletic clusters when applied to the Thai data set, characterized by persistent transmission alongside sparse and non-systematic sampling. Furthermore, we found that clusters derived using low distance thresholds could notably vary between the pwSNP-distance and comparative tree-based clustering methods, suggesting limited reliability and robustness. To accurately delineate tuberculosis transmission clusters, especially for complex data from high-burden, endemic settings, we recommend transitioning from pwSNP-distance clustering toward more robust, phylogenetic clustering that respects evolutionary descent.

Mycobacterium tuberculosis

The reverse DNA sequencing using Bst DNA polymerase.

The reverse DNA sequencing (RDS) [1] is a rapid method used to check the DNA sequences by sequencing them from the opposite orientation. Because the RDS is basically a double stranded sequencing, the quality of the sequence patterns so obtained generally is not as good as those obtained by the single stranded sequencing, and extra bands and higher background are produced more frequently. This paper shows that the RDS could now generate as good sequence patterns as those obtained by sequencing on the single stranded DNA template if Bst DNA polymerase instead of the conventional enzymes, such as the Klenow enzyme, was used in the RDS. Bst DNA polymerase is heat stable (optimum reaction temperature 65 degrees C) and has recently been successfully used in the conventional DNA sequencing. The RDS has recently been further simplified to meet the need of large DNA sequencing projects such as the human genome project. The combination of the simplified RDS and the use of Bst polymerase should be expected to facilitate greatly the work on sequence confirmation and correction.

Base Sequence

OligoSeq: Rapid nanopore-sequencing of single-stranded oligonucleotides.

Nanopore-based DNA sequencing technology has achieved remarkable success in sequencing increasingly long DNA strands (e.g., over a million nucleotides long) for genomics research and biotechnology applications. However, the same level of progress has not been achieved for DNA oligonucleotides (usually &#x2264; 300 nucleotides long). Oligonucleotides play a crucial role in genome engineering efforts through oligo library generation and in DNA data storage, where they are used to encode computer information, such as binary (digital) data in DNA libraries. To enable these applications, accurate sequencing of oligonucleotides in a way that allows to assess for sequence variability, quality and length is essential. But sequencing solutions for oligonucleotides - particularly DNA primers for PCR, oligo DNA libraries used for mutagenesis or cDNA libraries used in gene expression analysis - remain inadequate. To address this gap, OligoSeq is presented as an innovative approach that integrates two complementary techniques: AmpliSeq (based on PCR) and RevSeq (based on reverse complementation with sequence-specific or random primers) to facilitate sequencing of single-stranded oligonucleotides using reference sequence anchor matches of more than &#x2265; 90% identity spanning from about 70% to 10% with AmpliSeq or RevSeq with random nonamers, respectively, and resolving the final reference sequence based on the most likely candidate from basecall frequencies, regardless of length and double-stranding method. OligoSeq can be integrated with nanopore sequencing technology pipelines and can be used as a reference for other sequencing platforms requiring double-stranded adapters, offering a practical and scalable alternative for standard quality control in single-stranded oligonucleotide synthesis. The use of nanopore technology, compatible with the double-stranding methods showcased, is shown to be the most cost-effective method for resolving original DNA sequences of different length and quality, and to assess its sequence variability, compared to other methods such as Illumina, PacBio or HPLC/MS.

Sequence Analysis, DNA

Heat Inactivation of Nipah Virus for Downstream Single-Cell RNA Sequencing Does Not Interfere with Sample Quality.

Single-cell RNA sequencing (scRNA-seq) technologies are instrumental to improving our understanding of virus-host interactions in cell culture infection studies and complex biological systems because they allow separating the transcriptional signatures of infected versus non-infected bystander cells. A drawback of using biosafety level (BSL) 4 pathogens is that protocols are typically developed without consideration of virus inactivation during the procedure. To ensure complete inactivation of virus-containing samples for downstream analyses, an adaptation of the workflow is needed. Focusing on a commercially available microfluidic partitioning scRNA-seq platform to prepare samples for scRNA-seq, we tested various chemical and physical components of the platform for their ability to inactivate Nipah virus (NiV), a BSL-4 pathogen that belongs to the group of nonsegmented negative-sense RNA viruses. The only step of the standard protocol that led to NiV inactivation was a 5 min incubation at 85 &#xb0;C. To comply with the more stringent biosafety requirements for BSL-4-derived samples, we included an additional heat step after cDNA synthesis. This step alone was sufficient to inactivate NiV-containing samples, adding to the necessary inactivation redundancy. Importantly, the additional heat step did not affect sample quality or downstream scRNA-seq results.

Nipah Virus