Search PubMedSearch

Biomedical subjects

Tanya Golubchik

Publications and source records attributed to Tanya Golubchik.

5 recordsLinked to original sources

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans

Characterisation of Bordetella pertussis virulence and macrolide resistance in Australia by targeted culture-independent sequencing: a genomic epidemiology study.

BACKGROUND: Bordetella pertussis continues to circulate globally despite widespread vaccination, with a notable epidemic in 2024. Its resurgence is confounded by the emergence of pertactin-deficient, macrolide-resistant B pertussis strains in Asia and Europe, which are under-recognised by conventional diagnostics. We aimed to apply targeted culture-independent next-generation sequencing (tNGS) of respiratory specimens to improve global B pertussis diagnostic capability and genomic surveillance. METHODS: We did a nationwide genomic epidemiology study of B pertussis RT-PCR-positive respiratory specimens that were retrospectively and prospectively collected by diagnostic and public health laboratories in six of seven states and territories of Australia. Specimens underwent tNGS and macrolide-resistant B pertussis-specific PCR, and an opportunistic subset from New South Wales and Queensland were cultured for confirmatory susceptibility testing and whole-genome sequencing. Sequencing data were analysed for genome recovery, virulence profiles, and macrolide resistance mutations, and were compared with international macrolide-resistant B pertussis genomes and ancestral Australian genomes. The performance of the tNGS approach was assessed with logistic regression relative to RT-PCR cycle threshold values, and sensitivity and specificity values were calculated. FINDINGS: 255 respiratory specimens positive for B pertussis were included in the study. 64 (25%) were retrospectively collected between Jan 12, 2012, and Dec 31, 2023, and 191 (75%) were prospectively collected between Jan 1 and Oct 28, 2024. Of these 255 specimens, 148 (58%) yielded near-complete B pertussis genomes through tNGS. Seven co-circulating lineages of B pertussis were documented, including two associated with macrolide-resistance. Eight epidemiologically unrelated and geographically dispersed cases of macrolide-resistant B pertussis with a 23S rRNA 2037A→G mutation were identified by tNGS and confirmed by whole-genome sequencing. Three of these were further validated by phenotypic testing. The estimated prevalence of macrolide resistance among Australian cases positive for B pertussis was 4% (eight of 188). INTERPRETATION: tNGS can recover near-complete B pertussis genomes directly from clinical specimens, enabling identification of macrolide resistance mutations and high-resolution phylogenetic analysis. These findings show that tNGS complements PCR-based surveillance by providing genome-wide assessment of resistance, virulence, and genomic diversity in a single workflow. FUNDING: NSW Health Prevention Research Support Program.

Macrolides

Next-Generation Sequencing Methods for Sensitive Hepatitis B Viral Genome Analysis: A European Study.

This multicentre study investigated the utility of next-generation sequencing (NGS) to detect and generate hepatitis B virus (HBV) genomes in samples of low viral load (from 0.2 to 6207 IU/mL). 23 HBV DNA-positive plasma samples of genotypes A-E and one HBV-negative control sample were assayed blindly via 9 established NGS methods from 6 European laboratories. Methods included untargeted metagenomics, pre-enrichment by probe-capture followed by Illumina sequencing, and HBV-specific PCR pre-amplification followed by sequencing with Nanopore or Illumina. Full HBV genomes were obtained only from samples with viral loads > 1000 IU/mL using probe-capture methods, > 200 IU/mL using PCR-Illumina methods, > 10 IU/mL using PCR-Nanopore methods, and in no samples using metagenomic methods. Contamination was observed in the negative control and samples with very low viral loads in PCR-based methods. Probe-capture and metagenomic methods detected additional viruses not routinely screened in blood donations, including polyomaviruses and herpesviruses; positive results were confirmed by PCR. In conclusion, NGS may delineate whole-genome sequences at low viral loads if supported by a PCR pre-amplification step. Probe-capture methods also reliably detect HBV without pre-amplification but show limited genome coverage for samples with low viral loads; they may additionally detect a wide range of blood-borne viruses.

Humans

HIV-phyloTSI: subtype-independent estimation of time since HIV-1 infection for cross-sectional measures of population incidence using deep sequence data.

BACKGROUND: Estimating the time since HIV infection (TSI) at population level is essential for tracking changes in the global HIV epidemic. Most methods for determining TSI give a binary classification of infections as recent or non-recent within a window of several months, and cannot assess the cumulative impact of an intervention. RESULTS: We developed a Random Forest Regression model, HIV-phyloTSI, which combines measures of within-host diversity and divergence to generate continuous TSI estimates directly from viral deep-sequencing data, with no need for additional variables. HIV-phyloTSI provides a continuous measure of TSI up to 9 years, with a mean absolute error of less than 12 months overall and less than 5 months for infections with a TSI of up to a year. It performs equally well for all major HIV subtypes based on data from African and European cohorts. CONCLUSIONS: We demonstrate how HIV-phyloTSI can be used for incidence estimates on a population level.

HIV Infections

SARS-CoV-2 genomic diversity and within-host evolution in individuals with persistent infection in the UK: an observational, longitudinal, population-based surveillance study.

BACKGROUND: Persistent SARS-CoV-2 infections in hospitalised immunocompromised individuals are known to facilitate accelerated within-host viral evolution, potentially contributing to the emergence of highly divergent variants. However, little is known about the evolutionary dynamics and transmission risks of persistent infections in the general population. We aimed to characterise the within-host evolution of SARS-CoV-2 during persistent infections identified through a large community surveillance study. METHODS: We used data from the Office for National Statistics COVID-19 Infection Survey (ONS-CIS), a large-scale, longitudinal, population-based surveillance study conducted in the UK from April, 2020, to March, 2023. For this analysis, we focused on infections with high viral load (cycle threshold &#x2264;30) and available genome sequences, from seven major SARS-CoV-2 lineages (alpha, delta, BA.1, BA.2, BA.4, BA.5, and XBB). ONS-CIS participants were randomly selected from the general population and tested regularly by RT-PCR, regardless of symptoms. We defined persistent infections as those with sustained or rebounding high viral RNA titres for 26 days or longer. We examined associated host characteristics and used raw sequence data to identify de novo mutations and estimate within-host synonymous and non-synonymous evolutionary rates across the SARS-CoV-2 genome. FINDINGS: Between Nov 2, 2020, and March 21, 2023, we identified 576 persistent infections with at least two sequences, including 11 alpha, 106 delta, 102 BA.1, 204 BA.2, 16 BA.4, 133 BA.5, and 4 XBB. Persistent infections were more common in males than females (p<0&#xb7;0001) and individuals older than 60 years (p=0&#xb7;0027). The median within-host genome-wide evolutionary rate was 7&#xb7;9&#x2009;&#xd7;&#x2009;10-4 substitutions per site per year (IQR 7&#xb7;0-9&#xb7;0&#x2009;&#xd7;&#x2009;10-4), with high inter-individual variability driven largely by non-synonymous mutations, particularly in the N-terminal and receptor-binding domains of the spike protein. Longer infection duration was associated with higher evolutionary rates, while no associations were found with age, sex, vaccination status, previous infection, or virus lineage. We found no clear evidence of transmission beyond the first month of infection in any of the 84 persistent infections lasting 56 days or longer. In total, we identified 379 recurrent mutations, including many with known or predicted negative fitness effects and low prevalence at the population level, as well as de novo reversions to the Wuhan-Hu-1 reference sequence, which were likely under positive selection within those individuals. INTERPRETATION: This study highlights the heterogeneous nature of within-host SARS-CoV-2 evolution in individuals with persistent infection in the community. Notably, a small subset of persistent infections with high viral loads underwent accelerated viral evolution or recurrently acquired hallmark mutations found in novel variants. In addition, onward transmission from a persistent infection during the later stages of infection is likely to be rare. These insights have important implications for prioritising genomic surveillance and managing patients with persistent infections. FUNDING: Department of Health and Social Care.

Humans