Search PubMedSearch

SEARCH · Search PubMed

Results for “intrahost variants”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

5 recordsLinked to original sources

Using intrahost single nucleotide variant data to predict SARS-CoV-2 detection cycle threshold values.

Over the last four years, each successive wave of the COVID-19 pandemic has been caused by variants with mutations that improve the transmissibility of the virus. Despite this, we still lack tools for predicting clinically important features of the virus. In this study, we show that it is possible to predict the PCR cycle threshold (Ct) values from clinical detection assays using sequence data. Ct values often correspond with patient viral load and the epidemiological trajectory of the pandemic. Using a collection of 36,335 high quality genomes, we built models from SARS-CoV-2 intrahost single nucleotide variant (iSNV) data, computing XGBoost models from the frequencies of A, T, G, C, insertions, and deletions at each position relative to the Wuhan-Hu-1 reference genome. Our best model had an R2 of 0.604 [0.593-0.616, 95% confidence interval] and a Root Mean Square Error (RMSE) of 5.247 [5.156-5.337], demonstrating modest predictive power. Overall, we show that the results are stable relative to an external holdout set of genomes selected from SRA and are robust to patient status and the detection instruments that were used. This study highlights the importance of developing modeling strategies that can be applied to publicly available genome sequence data for use in disease prevention and control.

SARS-CoV-2

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans

Longitudinal analysis of high-risk HPV infections reveals within-host viral genome changes over time.

Persistent infection with high-risk (HR)-HPV causes cervical cancer, however, it is unclear why most infections resolve while a minority progress. We deep sequenced the HPV genomes of 1,228 HR-HPV-positive serial samples from 351 women with persistent infections (2-10 serial samples per woman over 1-8 years), including 279 controls and 72 precancer/cancer cases, to assess HR-HPV genome changes during infection and relation to infection outcomes. Seventy-seven percent of persistent infections (45-97% by HPV type) were infections with the same exact viral genome isolate; for HPV16, only 52% were persistent with the same isolate. This may suggest some infections include a type-specific isolate switch or new isolate infection during persistence. We additionally observed within-host change to the HPV genome estimated as gradual changes to intrahost single nucleotide variant (iSNV) frequency, and changes varied by HPV type, with HPV33 infections showing the most iSNV changes. Cases exhibited fewer viral genome changes during infection compared to controls (OR = 0.31, 95% CI = 0.1 - 0.86, p = 0.019), suggesting a more stable and clonal viral genome in cases. By viral gene, E7 had fewer nonsynonymous mutations in the cases compared to controls that cleared within 2 years of infection (p = 0.012), which confirms the importance of E7 conservation and suggests mutations to E7 reduce persistence associated with progression. There was a similar pattern in E4 (p = 0.013), while E5 had more changes in the cases (p = 0.008). A subset of 28 infections had an intervening HPV-negative sample between HPV-positive visits; 93% of these infections had the same exact viral genome isolate in the samples before and after the negative, consistent with subclinical persistence and subsequent re-detection. Our data suggests that HR-HPV type-persistence can include a collection of viral isolates, and viral mutations during infection, particularly in E7, reduce HR-HPV persistence and thus carcinogenic potential.

Humans

No receptor-binding domain adaptation detected in within-host H5N1 surveillance of 4,559 US dairy outbreak sequences.

BACKGROUND: The 2024-2026 US H5N1 clade 2.3.4.4b dairy cattle outbreak has been characterised primarily through consensus-level phylogenetics. Whether mammalian-adaptation variants are emerging at sub-consensus frequencies within infected hosts, particularly at the haemagglutinin receptor-binding domain (RBD), remains unknown because no systematic within-host variant analysis of the public sequencing corpus has been performed. METHODS: We conducted a pre-registered, corpus-wide intrahost single-nucleotide variant (iSNV) analysis of all publicly available H5N1 cattle, feline-spillover, and retail-milk sequences on the NCBI Sequence Read Archive (4559 samples across 7 BioProjects). A dual-caller concordance pipeline (iVar + LoFreq) with empirically determined allele frequency (AF) threshold (3%, set via four-criterion validation including synthetic spike-in controls) was applied to an 11-site Tier 1 mammalian-adaptation panel spanning the polymerase complex, haemagglutinin RBD, and accessory proteins. Within-host nucleotide diversity was compared across host categories. RESULTS: The HA RBD sites Q226L and G228S (H3 numbering) showed zero detections across >4300 adequately sequenced samples at all AF thresholds tested (1-5%), despite the pipeline detecting other non-synonymous variants at these exact codon positions (upper 95% CI for prevalence: 0.08%). Seven of eleven adaptation sites carried statistically significant iSNV signals after Bonferroni correction (corrected α = 0.00417), though all at low prevalence (≤2.95%). Genotype stratification showed that most polymerase-site detections reflected genotype structure rather than within-host emergence: the apparent PB2 631 L→M "reversion" was largely the ancestral avian state of the D1.1 genotype (20 of 23 detections), which never acquired the 631L mammalian adaptation, with only two genuine sub-consensus events in the B3.13 background, while consensus-level PB2 701N was a fixed feature of the D1.1 genotype (10 of 14 detections) rather than independent sub-consensus emergence. Cattle exhibited significantly higher within-host nucleotide diversity than feline-spillover samples (π = 1.59 × 10-4 vs 6.11 × 10-5; Kruskal-Wallis p = 6.6 × 10-15), a finding that persisted after depth-matching (p = 4.6 × 10-5); this may reflect prolonged mammary-gland infection, though sampling differences and host biology cannot be excluded. CONCLUSIONS: We did not detect HA receptor-switching adaptation (the acquisition of human-type α2,6 receptor binding via Q226L/G228S) at any tested allele frequency in the US dairy H5N1 outbreak. Sub-consensus mammalian-adaptation signals exist at polymerase-complex sites but at low prevalence, are genotype-structured rather than independently recurrent, and require functional characterisation before informing risk assessment.

Dairy cattle

Characterisation of a persistent SARS-CoV-2 infection lasting more than 750 days in a person living with HIV: a genomic analysis.

BACKGROUND: People who are immunocompromised can develop persistent SARS-CoV-2 infections. Several viral mutations accumulated during the course of such persistent infections have also been observed in prominent variants of concern (VOCs). Here, we characterise persistent infection and viral evolution of SARS-CoV-2 lasting more than 750 days in a person with advanced HIV-1 infection. METHODS: Between March, 2021, and July, 2022, eight clinical specimens were collected from a person living with HIV, neither receiving antiretroviral therapy nor virally suppressed, and presumed to have been initially infected with SARS-CoV-2 in mid-May, 2020. Viral RNA was extracted from each swab and an amplicon-based sequencing approach was used for genomic analysis of SARS-CoV-2. Variable sites were characterised at the consensus and subconsensus levels, and phylogenetic tools were applied to analyse viral evolution. Publicly available SARS-CoV-2 sequences from GenBank were leveraged to contextualise our sequenced samples and identify any potential evidence of transmission. FINDINGS: Genomes formed a monophyletic cluster in the B.1 lineage. 68 consensus and 67 subconsensus single nucleotide variants were observed over the course of infection. The intrahost clock rate remained similar to that of the interhost rate in contemporaneous community sequences (6·74 × 10-4 [95% credible interval 5·05 × 10-4 to 8·54 × 10-4] substitutions per site per year vs 6·11 × 10-4 [5·54 × 10-5 to 6·66 × 10-4]). Mutations grouped into two distinct subpopulations present throughout infection. 10 non-synonymous mutations in the spike protein gene were at positions in common with those defining the omicron lineage (BA.1 or BA.2), of which nine were present before November, 2021. Nine of 18 substitutions present throughout infection were rare in online databases, suggesting a lack of long transmission chains descending from this individual. INTERPRETATION: Convergent SARS-CoV-2 evolution, both in and outside the spike protein, observed in this study suggests parallels with the evolutionary process leading to emergence of the omicron VOC. The inferred absence of onward infections might indicate a loss of transmissibility during adaptation to a single host. Our results underscore the importance of appropriate treatment to cure persistent SARS-CoV-2 infections and monitoring them to understand how mutations contribute to viral adaptation. FUNDING: National Institute of General Medical Sciences of the National Institutes of Health, Centers for Disease Control and Prevention, the National Institute of Allergy and Infectious Diseases, MassCPR, and Morris Singer Foundation.

Humans