Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “single-nucleotide polymorphisms”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Determinants of functional burden pleiotropy and gene dosage responses across human traits.

Pleiotropic and monotonic effects of gene dosage are central to understanding comorbidities in developmental pediatric and psychiatric disorders, yet the underlying biological processes are not well characterized. Here we develop a functional burden analysis to investigate the association of all protein-coding copy-number variants, genome-wide, with 43 complex traits in approximately 500,000 UK Biobank participants. We test variant associations disrupting 172 tissue or cell-type gene sets, finding associations for all traits, which we replicate in the All of Us cohort. Functional burden pleiotropy, defined as the number of traits significantly associated with a gene set, correlates with genetic constraint and is higher for brain than non-brain functions, even after normalizing for genetic constraint. Levels of pleiotropy, measured by burden correlation, are similar in deletions and loss-of-function single-nucleotide variants, and higher than in common variants and duplications. Most gene dosage responses are non-monotonic, with deletions and duplications showing same-direction effects, and monotonic responses decrease with genetic constraint. We observe associations between functional gene sets and traits for either deletions or duplications, but rarely both, with negatively correlated effect sizes. Together, these results link genetic constraint and brain-specific mechanisms to the whole-body multimorbidity of neurodevelopmental and psychiatric conditions.

Humans↗

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans↗

Characterization of phosphorylation variants for identifying adaptive alleles in Zea.

Large-scale genome sequencing of maize wild species (teosinte) has uncovered thousands of genetic mutations, but distinguishing causal alleles from neutral variations remains a significant challenge. In this study, we conducted a comprehensive analysis of phosphorylation-associated single-nucleotide variations (pSNVs) to enhance our understanding of adaptive variations in the Zea genus. We collected 234 teosinte genomes from seven different taxa and 507 cultivated maize genomes to identify single-nucleotide variants that target phosphorylation machinery, which is crucial for plant development and environmental adaptation. Our analysis identified 33 687 pSNVs within the Zea genus and revealed a reduction in genetic conservation along with an increase in protein abundance and expression for genes harboring pSNVs. Additionally, pSNVs present stronger purifying selection pressures compared with other missense mutations. We found that maize possesses fewer pSNVs than teosinte, likely due to the effects of selection and hitchhiking. By examining the role of pSNVs related to kinase-substrate rewriting events and exhibiting evolutionary divergence jointly, our results suggest that pSNVs impact multiple traits, particularly flowering time variation between teosinte and maize. Furthermore, we documented the widespread presence of pSNVs in Arabidopsis thaliana, rice, and wheat, identifying 46 pSNVs that have convergently evolved between maize and other species. Our study provides another insight into uncovering adaptive alleles in wild species by incorporating protein signaling sites and emphasizes the potential of utilizing wild species for future crop improvement.

Zea mays↗

G4SNVHunter: An R/Bioconductor Package for Evaluating SNV-Induced Disruption of G-Quadruplex Structures Leveraging the G4Hunter Algorithm.

G-quadruplexes (G4s) are nucleic acid secondary structures with important regulatory functions. Single-nucleotide variants (SNVs), one of the most common forms of genetic variation, can potentially impact the formation of G4 structures if they occur within G4 regions. However, there is currently a lack of software tools specifically designed to assess such effects. Here, we present an R/Bioconductor package named G4SNVHunter, which enables rapid detection of variants that may disrupt G4 structures. This tool, based on the core principles of the G4Hunter algorithm, can provide precise quantitative assessment of the propensity for G4 formation within genomic sequences. Specialized experimental methods can then be designed based on the results provided by G4SNVHunter to further verify the specific functions of the affected G4 structures, facilitating deeper insights into the biological impacts of genetic variants from the perspective of G4 structures. To showcase the functionality of the G4SNVHunter package, we analyzed the Neandertal and Denisovan archaic introgressed variants detected by the Sprime software, and identified approximately 5,800 variants located within G4 regions, among which around 230 may impair G4 structure formation propensity. The source code for the G4SNVHunter package has been publicly released under the MIT license at https://github.com/rongxinzh/G4SNVHunter and https://bioconductor.org/packages/devel/bioc/html/G4SNVHunter.html.

G-Quadruplexes↗

Plasma lipid species, immune cell traits, and gastric cancer risk: A Mendelian randomization study.

Plasma lipid composition has been linked to multiple cancers, yet its causal contribution to gastric cancer and the potential intermediary role of immune cells remain unclear. We aimed to clarify these relationships and identify specific lipid and immune cell traits that either protect against or promote gastric cancer. We performed a 2-sample, 2-step Mendelian randomization analysis using summary statistics from large genome-wide association studies of gastric cancer (1423 cases, 3,14,193 controls), plasma lipidomics (179 molecular species), and 731 immune cell phenotypes. Independent, genome-wide significant single-nucleotide variants served as instrumental variables. First, we estimated the causal effects of each plasma lipid on gastric cancer. Then, we explored the potential intermediary role of lipid-associated immune cell traits using a 2-step Mendelian randomization framework. Two lipids - phosphatidylethanolamine (18:0 0:0) and phosphatidylcholine (O-18:0 16:1) - were causally associated with a lower risk of gastric cancer. Three immune cell traits (CD8br and CD8dim %leukocyte, IgD on IgD+ CD38- and CD3 on CD28- CD8br) similarly showed protective effects. In contrast, phosphatidylcholine (O-16:1 18:2), triacylglycerol (49:1), triacylglycerol (56:3), and triacylglycerol (56:4) increased gastric cancer risk, as did immune traits such as TD DN (CD4-CD8-)AC, CD19 on memory B cell, CD28 on CD39+ activated Treg, CD45 on CD4+, CD127 on CD28+DN(CD4-CD8-) and CCR2 on CD14+CD16+ monocyte. Exploratory mediation analyses found no statistically significant evidence that immune cell traits mediated the effects of plasma lipids on gastric cancer risk. Specific phosphatidylethanolamines and phosphatidylcholines confer protection against gastric cancer, whereas several triacylglycerols increase risk. However, exploratory mediation analyses provided no statistically significant evidence that immune cell traits mediated these associations.

Humans↗

High-accuracy SNV calling for bacterial isolates using deep learning with AccuSNV.

Accurate detection of mutations within bacterial species is critical for fundamental studies of microbial evolution, reconstruction of transmission events, and identification of antimicrobial resistance mutations. Although many tools have been developed to identify single-nucleotide variants (SNVs) from whole-genome sequencing, they often suffer from high false-positive rates owing to the complexity of bacterial genomes and the need for different filtering cutoffs across sample types and sequencing depths. As data sets increase in size, the manual filtering required for high accuracy presents a significant obstacle. Here, we present AccuSNV, a novel deep learning-based tool for high-precision and automated bacterial SNV calling. Unlike traditional methods that process one sample at a time, AccuSNV leverages a convolutional neural network (CNN) that integrates alignment information across multiple samples, enhancing precision through learned across-sample patterns. We evaluate AccuSNV against seven popular SNV-calling tools using simulated data from six bacterial species with varied sequencing depths, numbers of isolates, mutations, and divergence levels. To further validate its real-world utility, we test AccuSNV on multiple curated bacterial data sets containing reported SNVs. In both simulated and real-world scenarios, AccuSNV consistently achieves the best performance. Moreover, AccuSNV provides comprehensive user-friendly downstream analysis modules and outputs, including mutation annotation information, phylogenetic inference, d N/d S calculations, and optional manual filtering. Together with the automated deep learning-based calling, these features make AccuSNV broadly accessible to users with different levels of computational expertise.

Deep Learning↗

Visual Detection and Stratification of Pathogenic mtDNA SNV Heteroplasmy by Balancing FnCas12a Signal Output and Allelic Discrimination.

Assessment of pathogenic mitochondrial DNA (mtDNA) single-nucleotide variant (SNV) heteroplasmy is important for molecular diagnostics, yet rapid visual profiling remains analytically challenging because an assay must combine single-nucleotide allelic discrimination, mutant-fraction-associated readout, and suitable target access. Herein, we report VISTA (visual identification and stratification of targeted mtDNA alleles), a broad-PAM FnCas12a assay that rebalances trans-cleavage signal output and mutant-wild-type discrimination for visual mtDNA SNV heteroplasmy analysis. VISTA uses unmodified FnCas12a with relaxed TTN PAM recognition and integrates crRNA spacer-length engineering with PEG8000/acBSA reaction tuning to improve the practical signal-discrimination balance without nuclease engineering. At the m.3243A>G model locus, spacer truncation enhanced mutant-wild-type discrimination, while molecular-dynamics simulations identified spacer-dependent differences between matched and mismatched complexes at the crRNA-DNA interface. The optimized assay resolved defined synthetic m.3243A>G heteroplasmy gradients by fluorescence imaging and was further adapted to lateral-flow detection. In locus-specific analyses of a deidentified collection of 74 peripheral-blood samples, fluorescence and lateral-flow readouts achieved ROC AUC values above 0.9 for mutant-allele classification after target-region amplification. Fluorescence supported heteroplasmy-associated profiling, whereas lateral flow provided a visual, semiquantitative readout for relative ranking based on the T/C ratio rather than absolute heteroplasmy measurement. VISTA therefore provides an accessible dual-readout analytical strategy for visual detection and heteroplasmy-associated profiling by tuning the FnCas12a signal output and allelic discrimination.

DNA, Mitochondrial↗

Mechanism of age-related accumulation of mtDNA mutations in human blood.

Accumulation of mutant mitochondrial DNA (mtDNA) heteroplasmy is among the strongest signatures of ageing1. Here we investigated the underlying mechanism by calling mtDNA sequence, mtDNA abundance and mtDNA heteroplasmic variants in human blood using whole-genome sequences from approximately 750,000 individuals. We observed that mtDNA single-nucleotide variants (mtSNVs) accumulate sharply at age 60 years, occur at low levels of heteroplasmy, exhibit little evidence of positive selection and are likely to be predominantly neutral. The mutational spectrum of mtSNVs does not reflect oxidative lesions, as is commonly invoked, but is more consistent with mtDNA replication errors. To understand why mtSNVs become detectable with age, we performed a genome-wide association study for heteroplasmic mtSNV burden, identifying germline variants near TERT, TCL1A and SMC4, all of which have been linked to clonal haematopoiesis (CH)2. Rare-variant analysis also showed that high mtSNV burden is associated with mutations in numerous CH driver genes. These genetic associations persisted even after exclusion of individuals with known CH driver mutations. Our results support a model in which 'cryptic' mtDNA mutations initially arise randomly as replication errors but are undetectable in bulk. They then become apparent only through age-related expansion of cellular clones in blood. We propose that the high copy number and mutation rate of mtDNA make it a sensitive blood-based marker of somatic mosaicism due to CH. Our work mechanistically unifies three prominent signatures of ageing: common germline variants in TERT, CH and observed accrual of mtDNA mutations.

Humans↗

Cross-kingdom genomic variation in chicken gut microbiomes: insights from China's diverse local breeds.

BACKGROUND: The gut microbiome possesses substantial genetic diversity that supports microbial adaptation, but the genomic variation patterns across its prokaryotic and viral populations remain incompletely characterized. RESULTS: Through integrated metagenomic and metatranscriptomic analysis of ten indigenous chicken breeds from China, we recovered 1527 representative prokaryotic MAGs, 37,555 representative DNA viral contigs, and 1867 representative RNA viral contigs (primarily comprising Bacillota/Bacteroidota, Uroviricota, and Lenarviricota/Pisuviricota, respectively). By integrating complementary short-read and long-read metagenomics with metatranscriptomics, we identified structural variants (SVs) and single-nucleotide variants (SNVs) in these cross-kingdom genomes. Positive SV-SNV density correlations occurred consistently across all microbial groups, indicating coordinated mutational processes. DNA viruses exhibited the highest variant prevalence (86.9% SNVs, 47.7% SVs), with temperate phages accumulating significantly more variants than virulent phages. Functionally, prokaryotic variants accumulated in carbohydrate metabolism and amino acid metabolism, while viral variants demonstrated broad metabolic hijacking. Horizontal gene transfer (HGT) was characterized by a strong virus-associated signature (69.40% of 536 events) and marked by an asymmetric pattern, with phage-to-bacteria (P-to-B) flow alone constituting 37.50% of all events. Random forest analysis revealed a strong bidirectional predictive relationship between SV and SNV densities across prokaryotic, DNA viral, and RNA viral populations, suggesting coupled genomic instability. Niche breadth emerged as a major driver of SNVs across kingdoms and was positively correlated with variant density. In prokaryotes, HGT events significantly shaped variant patterns. For viruses, genomic GC content was an important factor and consistently showed a negative correlation with SNV density in both DNA and RNA viruses. CONCLUSIONS: These findings demonstrate that coordinated mutational processes and kingdom-specific intrinsic factors drive genomic variation, with viruses serving as key genetic exchange vectors in chicken gut ecosystems. Video Abstract.

Animals↗

Novel Genetic Loci in Early-Onset Gout Derived From Whole-Genome Sequencing of an Adolescent Gout Cohort.

OBJECTIVE: Mechanisms underlying the adolescent-onset and early-onset gout are unclear. This study aimed to discover variants associated with early-onset gout. METHODS: We conducted whole-genome sequencing in a discovery adolescent-onset gout cohort of 905 individuals (gout onset 12 to 19 years) to discover common and low-frequency single-nucleotide variants (SNVs) associated with gout. Candidate common SNVs were genotyped in an early-onset gout cohort of 2,834 individuals (gout onset &#x2264;30 years old), and meta-analysis was performed with the discovery and replication cohorts to identify loci associated with early-onset gout. Transcriptome and epigenomic analyses, quantitative real-time polymerase chain reaction and RNA sequencing in human peripheral blood leukocytes, and knock-down experiments in human THP-1 macrophage cells investigated the regulation and function of candidate gene RCOR1. RESULTS: In addition to ABCG2, a urate transporter previously linked to pediatric-onset and early-onset gout, we identified two novel loci (Pmeta < 5.0 &#xd7; 10-8): rs12887440 (RCOR1) and rs35213808 (FSTL5-MIR4454). Additionally, we found associations at ABCG2 and SLC22A12 that were driven by low-frequency SNVs. SNVs in RCOR1 were linked to elevated blood leukocyte messenger RNA levels. THP-1 macrophage culture studies revealed the potential of decreased RCOR1 to suppress gouty inflammation. CONCLUSION: This is the first comprehensive genetic characterization of adolescent-onset gout. The identified risk loci of early-onset gout mediate inflammatory responsiveness to crystals that could mediate gouty arthritis. This study will contribute to risk prediction and therapeutic interventions to prevent adolescent-onset gout.

Humans↗

AUTS2-related syndrome: Insights from a large European cohort.

PURPOSE: AUTS2-related syndrome is characterized by developmental delay, autism spectrum disorder, and intellectual disability. From alternative promoters, AUTS2 encodes 2 distinct long and short isoforms encoding a putative transcriptional activator. METHODS: Through a European collaborative study, we collected clinical and genotype data on the largest AUTS2-related syndrome cohort of 58 patients harboring genomic rearrangements or single-nucleotide variants (SNVs). RESULTS: Pathogenic SNVs were recurrently found in individuals from different countries, suggesting mutational hotspots. Independent of the underlying defect at the AUTS2 locus, we observed that autistic behavior, hyperactivity, learning difficulties, and speech delay are common features of AUTS2-related syndrome. Among patients with SNVs, individuals carrying pathogenic variants affecting both longer and shorter AUTS2 transcripts showed a recognizable phenotype with microcephaly, brachycephaly, microretrognathia, broad nasal base, and anteverted nares. Behavioral disorders were more common in patients with variants affecting only the longer isoform. Arthrogryposis and stiff movements were only observed in patients with SNVs. CONCLUSION: This study provides a comprehensive clinical characterization of AUTS2-related syndrome, reveals few genotype-phenotype correlations, and suggests that the disruption of the 2 distinct AUTS2 transcripts has a different impact on the clinical phenotype.

Humans↗

Deep clinical and genetic analysis of 17p13.3 region: 38 pediatric patients diagnosed using next-generation sequencing and literature review.

BACKGROUND: Chromosome 17p13.3 is a region of genomic instability associated with different neurodevelopmental diseases. The malformation spectrum of 17p13.3 microdeletions ranges from an isolated lissencephaly sequence to Miller-Dieker syndrome, while 17p13.3 microduplications result in autism, learning disabilities, microcephaly and other brain malformations. This study aims to provide a more comprehensive delineation of the clinical and genetic characteristics associated with 17p13.3 alterations. METHODS: We retrospectively analyzed the next-generation sequencing (NGS) data of more than 40 thousand patients from January 2016 to December 2021 and identified 38 pediatric patients with copy-number variations (CNVs) or single-nucleotide variations (SNVs) in 17p13.3 region. Published patients with CNVs in the 17p13.3 region were also collected and we performed a Chi-square test to compare the phenotype spectrum of microdeletions and microduplications. RESULTS: Among the 27 CNV patients, 20 patients with microdeletions and 7 patients with microduplications were found. PAFAH1B1 was the most frequently deleted gene and CRK was the most frequently duplicated gene. Affected genes in 11 SNV patients included PAFAH1B1 and PRPF8. Developmental delay was the most common abnormality detected in the 38 patients (29/38, 76.3%). Of note, Case 10 presented omphalocele and Case 23 presented scoliosis, webbed neck and bone cyst, all of which were unusual variant phenotypes in this region. The Chi-square test revealed that epilepsy, lissencephaly and short stature were statistically significant with microdeletions, while behavioral abnormalities and hand and foot abnormalities were significant with microduplications (p&#x2009;<&#x2009;0.01). CONCLUSIONS: While PAFAH1B1, YWHAE and CRK are associated with major phenotypes of 17p13.3, RTN4RL1 may be involved in white matter changes and HIC1 might contribute to the occurrence of omphalocele. This study provided a comprehensive understanding of genetic information and phenotype spectrum of the 17p13.3 region.

Humans↗

Combined somatic mutation and transcriptome analysis reveals region-specific differences in clonal architecture in human cortex.

The human cerebral cortex is specialized into regions, but little is known about how human cellular lineages shape cortical regional variation and neuronal cell-type distribution during development. Here, we map single-cell lineages of human cortical regions and neuronal subtypes using >1,000 somatic single-nucleotide variants (sSNVs) identified from deep bulk whole-genome sequencing and analyzed over 25 regions and >72,000 single cells. In the fronto-parietal cortex, sSNVs are rarely restricted, marking neuron-generating clones that disperse into neighboring regions. In contrast, the primary visual cortex harbors 30%-70% more sSNVs than the neighboring secondary visual cortex. Clones at this border exhibit more restricted dispersion, suggesting late developmental lineage segregation. Single-nucleus sSNV and whole-transcriptome analysis reveal glutamatergic neuron clones with modest regional restrictions that share low-mosaic sSNVs with some GABAergic neurons, suggesting a recent dorsal cortical progenitor. Our analysis reveals human-specific cortical lineage patterns, regional differences in clonal patterns, and late divergence of some glutamatergic/GABAergic lineages.

Humans↗

SeqQC-former: A sequence-quality fusion framework for QC-aware review prioritization of candidate somatic SNVs in cancer genomics.

The accurate prioritization of candidate somatic single-nucleotide variants (SNVs) remains a challenge due to the substantial variability in sequencing quality across genomic loci. SeqQC-Former is a sequence-quality fusion framework that integrates the local nucleotide context with read-level quality-control (QC) covariates derived from matched tumor-normal sequencing data. This integration generates QC-aware prioritization scores for the downstream review of candidate variants. Unlike conventional variant callers, SeqQC-Former is designed not to infer biological truth but to support post-calling review and prioritization under heterogeneous sequencing conditions. The framework was trained and evaluated on a SEQC2-derived dataset comprising 89,447 candidate loci, including 1378 positive and 88,069 negative loci. In chromosome-held-out validation, which aims to reduce potential genomic-position leakage, SeqQC-Former demonstrated strong discrimination (AUROC = 0.9479; AUPRC = 0.9448), indicating good generalization to previously unseen chromosomes. Given that the SEQC2-derived labels contain QC-associated information; these results should be interpreted as an evaluation of QC-aware prioritization capability rather than an independent validation of biological variant correctness. Ablation analyses revealed that structured QC covariates provided the dominant predictive signal under the current SEQC2-derived labeling regime. SeqQC-Former achieved a significantly higher AUROC than classical machine-learning baselines, as determined by DeLong's test (p&#x202f;<&#x202f;0.01). Application to 53,164 glioblastoma variants demonstrated that external predictions were sensitive to QC scaling and threshold selection, underscoring that model outputs should be interpreted as QC-dependent prioritization scores rather than calibrated probabilities or definitive biological classifications. Overall, SeqQC-Former offers a reproducible post-calling QC-aware prioritization framework for large-scale somatic SNV review and underscores the importance of explicitly modeling sequencing-quality information when interpreting structured cancer genomics datasets.

Humans↗

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans↗

scSNViz: visualization and analysis of cell-specific expressed SNVs.

MOTIVATION: Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS: scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION: scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.

Software↗

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans↗