Search PubMedSearch

SEARCH · Search PubMed

Results for “variational inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Accounting for recombination rate variation improves inference of barrier loci and reveals the role of both natural and sexual selection in an incipient bird radiation.

Examining genomic patterns of differentiation across lineage pairs at different stages of the speciation continuum, in combination with recombination maps, can help disentangle the effects of linked and divergent selection and identify lineage-specific targets of selection that may act as barrier loci during speciation. Here, we apply this framework to genomic data from African and Indian Ocean bird species of the genus Zosterops (Zosteropidae) to identify candidate barrier loci between ecologically, phenotypically, and genetically distinct Reunion gray white-eye (Zosterops borbonicus) parapatric geographic forms. Using analyses that account for recombination rate variation, we show that putative targets of divergent selection are primarily located on the Z chromosome, except in comparisons between geographic forms that differ in their ecologies. Functional annotation revealed that candidate barrier loci between forms with similar environmental niches are associated with genes involved in song formation and immune function, whereas those between forms with different environmental niches are associated with adaptation to altitude, morphology, and song behavior. Our results highlight the combined roles of natural and sexual selection in the evolution of reproductive barriers in this incipient species radiation.

Animals

Unifying multimodal single-cell data with a mixture-of-experts β-variational autoencoder framework.

Multimodal single-cell assays profile complementary layers of cell state, but integration is complicated by modality mismatch, sparsity, and uneven cohort coverage. Here, we present Unified Variational Inference (UniVI), a scalable mixture-of-experts β-variational autoencoder that learns a shared latent space while preserving modality-specific structure. UniVI couples modality-specific encoders/decoders with a shared latent prior and a symmetric cross-modal alignment objective, enabling consistent integration of paired measurements without curated feature-link graphs or preannotated reference atlases; optional supervised heads can be added when labels are available. Across paired RNA-protein (CITE-seq) and RNA-chromatin (10x Genomics Multiome, SHARE-seq) data spanning human PBMCs and mouse back skin-a nonhematopoietic tissue with continuous differentiation hierarchies-UniVI produces coherent embeddings, improves label transfer, and enables cross-modal reconstruction and denoising. Extending to trimodal measurements, UniVI maintains robust three-way alignment among RNA, chromatin accessibility, and surface proteins (TEA-seq), and accommodates DNA methylation in a paired scNMT-seq mouse gastrulation proof-of-concept under beta-binomial likelihoods. Performance degrades gracefully under severe cell type imbalance and in the presence of modality-exclusive populations. In an acute myeloid leukemia mosaic design, a paired RNA-protein bridge anchors independent RNA-only and protein+genotype cohorts, revealing genotype-associated neighborhoods that sharpen with mutation-aware fine-tuning. UniVI thus provides a flexible, interpretable framework for multimodal integration across paired, trimodal, and mosaic study designs and supports practical reference-to-query projection in partially observed studies.

Journal Article

Interactions between diverse proteinoids and microspheres in simulation of primordial evolution.

Experiments demonstrating an incorporation of different enzymelike activities into a single preparation of proteinoid microspheres provide a conceptual basis for the primitive lengthening of protometabolic pathways. An enhancement of one enzymelike activity by another proteinoid in the same microsphere has been found. This effect, plus the pathway-lengthening propensity of combinations of microspheres, indicates selective advantages contributing to adaptive protoselection. Data reported in this paper also bring into purview the concept of internally controlled variation. Inferences are derived for the origin of protosexuality in protocells. When allowance is made for a closer relationship to the environment than that needed in contemporary selection, the fundamental mechanistic requirements of protoevolution are regarded as met by the proteinoid microsphere.

Catalysis

Exploiting Omic Data to Advance Predictive Ecotoxicology.

Predicting species-specific chemical sensitivity using in silico approaches has the potential to transform environmental risk assessment, conservation, and biomonitoring, while reducing, and ultimately replacing, animal testing. Genomic and transcriptomic data capture extensive sensitivity-relevant variation, including differences in molecular targets, xenobiotic metabolism, and damage mitigation pathways. Large-scale sequencing initiatives therefore offer an unprecedented opportunity to address ecotoxicology's "too many species" problem. Although existing omic-based predictive tools provide proof of concept, they have so far been applied to a narrow set of relatively straightforward prediction scenarios. To achieve broader applicability, current and future tools must be firmly grounded in the diverse molecular mechanisms underlying differential chemical responses. Here, we critically evaluate the emerging field of predicting species sensitivity using molecular variation inferred from omic data. We analyze the strengths and limitations of current omic-based approaches and identify major sequence and ecotoxicological data gaps, as well as critical bioinformatic challenges. We then review the current knowledge of how molecular biology underlies differential chemical sensitivity, outlining research paths to allow the next generation of sensitivity prediction tools to exploit ever expanding omic data.

Ecotoxicology

NExON-Bayes: a Bayesian approach to network estimation informed by ordinal covariates.

MOTIVATION: In heterogeneous disease settings, accounting for intrinsic sample variability is crucial for obtaining reliable and interpretable omic network estimates. However, most graphical model analyses of biomedical data assume homogeneous conditional dependence structures, potentially leading to misleading conclusions. To address this, we propose a joint Gaussian graphical model that leverages sample-level ordinal covariates (e.g. disease stage) to account for heterogeneity and improve the estimation of partial correlation structures. RESULTS: Our modelling framework, called NExON-Bayes, extends the graphical spike-and-slab framework to account for ordinal covariates, jointly estimating their relevance to the graph structure and leveraging them to improve the accuracy of network estimation. To scale to high-dimensional omic settings, we develop an efficient variational inference algorithm tailored to our model. Through simulations, we demonstrate that our method outperforms the vanilla graphical spike-and-slab (with no covariate information), as well as other state-of-the-art network approaches which exploit covariate information. Applying our method to reverse phase protein array data from patients diagnosed with stage I, II or III breast carcinoma, we estimate the behaviour of proteomic networks as cancer progresses. Our model provides insights not only through inspection of the estimated proteomic networks, but also of the estimated ordinal covariate dependencies of key groups of proteins within those networks, offering a comprehensive understanding of how biological pathways shift across disease stages. AVAILABILITY AND IMPLEMENTATION: A user-friendly R package for NExON-Bayes with tutorials is available on Github at github.com/jf687/NExON, and archived at https://doi.org/10.5281/zenodo.20312938. The source of the dataset used is cited in the relevant section.

Bayes Theorem

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations

TreeFlow: Probabilistic Modelling and Automatic Differentiation for Phylogenetics.

Probabilistic modelling frameworks are powerful tools for statistical modelling and inference. They are not immediately generalizable to phylogenetic problems due to the particular computational properties of the phylogenetic tree object. TreeFlow is a software library for probabilistic modelling and automatic differentiation with phylogenetic trees. It embeds phylogenetic trees in the TensorFlow Probability framework, and implements inference algorithms for phylogenetic models given a fixed tree topology. We demonstrate how TreeFlow can be used to quickly implement and assess new models. We also show that it provides reasonable performance for gradient-based inference algorithms compared to specialized computational libraries for phylogenetics.

Bayesian inference

[Change in the event-related skin conductivity: an indicator of the immediate importance of elaborate information processing?].

In recent psychophysiological conceptualizations of the orienting response (OR) within the framework of information processing, the OR is increasingly considered a "call for processing resources", something which is especially inferred from variations in the event-related skin conductance response (SCR). The present study, therefore, was concerned with certain implications arising from this framework or perspective, particularly in regard to the question of whether stimuli eliciting skin conductance responses obligatorily receive/evoke processing priority or not. In order to examine whether these electrodermal responses denote a capturing of attention or merely a call for processing resources, short (1 s) pure sine tones of 65 dB with sudden onset (commonly used as orienting stimuli) were inserted in a reaction time paradigm with an additional memory load. This demand was primarily given because memory processes play a key role in theories of orienting and habituation. The task was run under two different conditions of complexity, factorially combined with a novelty variation of the added auditory stimuli. The results revealed a substantial deterioration of task performance subsequent to the occurrence of the tones, which, however, was dependent on task complexity and on novelty of the tones. The task impairment is particularly remarkable as subjects were asked to avoid distractions by paying attention to the task and as the tones were introduced as subsidiary and task-irrelevant. Together with the missing effects of task complexity on phasic and tonic electrodermal activity, results suggest that information-processing conceptualizations of the OR can only be a meaningful heuristic contribution to theoretical developments about human orienting and its habituation if the setting of processing priority, its conditions, as well as its implications are adequately taken into account. In addition, it seems to be promising to consider the strength of the SCR as an index of urgency of elaborate, attention-demanding processing and not as a peripheral physiological manifestation of the OR, or, respectively, of a call for unspecific processing resources. Such a view would also do justice to the aspect of prioritization. The sufficient conditions for an OR's occurrence could, in this context, be equated with, among others, some of those which activate a mechanism subserving selective attention and, as a possible result, which lead to further and more elaborate processing of potentially important information.

Adult

Gene-level complexity explains genome-wide variation in the distribution of fitness effects.

The distribution of fitness effects (DFE)-describing how harmful, neutral, or beneficial new mutations are-is central to understanding how populations evolve. Although the DFE varies across genomes and species, it remains unclear which aspects of genomic organization drive this variation. Here, we inferred gene-level selective constraints across the genomes of Mus musculus castaneus, Drosophila melanogaster and Saccharomyces cerevisiae using a combination of population genetics and machine learning trained on diverse gene features. Many gene features were predictive of selective constraint, with conservation, gene structure, and expression being the most informative. These selective constraints delineated gene classes with distinct DFEs. Genes with higher connectivity and expression-features reflecting how many traits a gene influences-experienced stronger and less dispersed deleterious effects with increasing selective constraint. Between species, the rate of adaptation decreased with increasing organismal complexity, whereas across the genome it did not decrease monotonically with selective constraint, but tended to be higher at intermediate levels. While between-species comparisons of DFE parameters were less consistent with predictions of Fisher's geometric model (FGM) based on organismal complexity, variation in DFE parameters across the genome aligned more closely with FGM when complexity was considered at the gene level. Our results suggest that gene-level complexity, captured by genomic feature proxies, provides a more informative definition of complexity for DFE variation than organism-level labels, and highlight the value of using gene features collectively to link genomic architecture, fitness landscapes, and patterns of molecular evolution.

Animals

A new aspect of the morphological transformation of Trypanosoma cruzi brought about by environmental variation.

The development of Trypanosoma cruzi has been described both in vertebrate and invertebrate hosts, and the morphological transformations of the parasite have been studied in both cell-free cultures and tissue cultures. The investigators who studied this topic have emphasized the fact that the morphogenesis of T. cruzi may be associated with a series of factors. In the present study, we noted that when bloodstream trypomastigotes leave a vertebrate host reaching the digestive tract of triatomines through the blood sucking action of these vectors, specific culture by blood plating or maintenance of blood in physiological saline at different temperatures shows a phenomenon of trypanosome joining, with intensive movement of internal organelles (nucleus and kinetoplast) and junction at the kinetoplast level. Different situations may occur after this phenomenon, such as flagellate separation, passage of kinetoplast content from one individual to another, transformation into rounded elements that approach the pairs of agglomerate, or the formation of spherical elements similar to cyst-like bodies. When observed by light or phase-contrast microscopy, these bodies appear to be static and show inner structures moving in circles or in disorderly manner. On the basis of the molecular studies carried out by other authors, who observed that not all proteins synthetized from DNA are of immediate usefulness in the cell, but need to undergo activation by the action of another protein or of environmental variation, we may infer that T. cruzi, under adverse conditions, i.e. a change in habitat, may undergo transformations, taking on different forms for the exchange of genetic information for adaptation to the environment and for possible continuity of the evolutionary cycle.

Animals

Circadian variation in the time of request for helicopter transport of cardiac patients.

STUDY OBJECTIVES: The literature has demonstrated circadian rhythms in the occurrence of nonfatal myocardial infarction, ischemia, and sudden death. We hypothesized that requests for helicopter transport of acutely ill cardiac patients followed a similar circadian pattern and differed significantly from requests for helicopter transport of other categories of patients. DESIGN: Prospective study of requests for helicopter transport of 1,128 consecutive air medically transported patients over a 24-month period. SETTING: One tertiary-care teaching hospital. MEASUREMENTS AND MAIN RESULTS: The periodic structure of the time distribution of cardiac requests for helicopter transport was examined with a two-harmonic regression analysis using a 24-hour period of oscillation. Seven hundred eighty-seven cardiac and 315 noncardiac patients could be evaluated. The times of requests for helicopter transport were tabulated into hourly intervals. Cardiac-related requests for helicopter transport were significantly different from noncardiac-related requests for helicopter transports (P less than .009 by Wilcoxon rank sum test, P less than .032 by Kolmogorov-Smirnov test). The regression model for cardiac requests for helicopter transport was also significant (P less than .0001, R2 = .81) with increasing requests for helicopter transport from 6:00 AM until 12:00 noon. CONCLUSION: The time distribution of requests for helicopter transport for cardiac patients demonstrates a striking circadian variation not observed in noncardiac patients. This observation strengthens mechanistic inferences from studies of circadian variation and suggests a "morning-loaded" staffing pattern for air medical services predominantly transporting cardiac patients.

Aircraft

Further issues in small area variations analysis.

In this article, I examine Wennberg's "practice style" hypothesis and the literature on variations among small areas. According to Wennberg, geographic variations in rates of per capita use for many clinical procedures arise mainly from differences over what constitutes appropriate care. I show, however, that the role of practice style in explaining variations among areas has not been clearly demonstrated. I also argue that the practice style hypothesis can neither be established nor refuted by the methods traditionally used to study small areas and, moreover, that inferences about practice style variations cannot be drawn from differences among areas in their rates of use. I thus conclude that more research at the micro level in the practice patterns of the individual physician is needed before major health care initiatives based on small area methodology are undertaken.

Health Policy

Sexual dimorphism in the human bony pelvis, with a consideration of the Neandertal pelvis from Kebara Cave, Israel.

Sexual dimorphism of the human pelvis is inferentially related to obstetrics. However, researchers disagree in the identification and obstetric significance of pelvic dimorphisms. This study addresses three issues. First, common patterns in dimorphism are identified by analysis of pelvimetrics from six independent samples (Whites and Blacks of known sex and four Amerindian samples of unknown sex). Second, an hypothesis is tested that the index of pelvic dimorphism (female mean x 100/male mean) is inversely related to pelvic variability. Third, the pelvic dimensions of the Neandertal male from Kebara cave, Israel are compared with those of the males in this study. The results show that the pelvic inlet is the plane of least dimorphism in humans. The reason that reports often differ in the identification of dimorphisms for this pelvic plane is that both the length of the pubis and the shape of the inlet are related to nutrition. The dimensions of the pelvis that are most dimorphic (that is, female larger than male) are the measures of posterior space, angulation of sacrum, biischial breadth, and subpubic angle. Interestingly, these dimensions are also the most variable. The hypothesis that variability and dimorphism are inversely related fails to be supported. The factors that influence pelvic variability are discussed. The Kebara 2 pelvis has a spacious inlet and a confined outlet relative to modern males, though the circumferences of both planes in the Neandertal are within the range of variation of modern males. The inference is that outlet circumference in Neandertal females is also small in size, but within the range of variation of modern females. Arguments that Neandertal newborns were larger in size than those of modern humans necessarily imply that birth was more difficult in Neandertals.

Animals

Spatial proteomics reveals four-stage molecular evolution in cancer immunotherapy-related gastritis.

BACKGROUND: Immune checkpoint inhibitors (ICIs) have transformed cancer treatment, yet immune-related adverse events (irAEs) including immunotherapy-related gastritis (IRAEG) pose significant clinical challenges-often necessitating treatment interruption that may compromise antitumor efficacy. IRAEG presents with atypical symptoms, lacks specific biomarkers, and shows histopathological overlap with other forms of gastritis, complicating diagnosis and management. Despite increasing clinical recognition, a systematic understanding of spatial molecular alterations across the full disease course remains limited. Here, we used spatial proteomics to map the molecular landscape of IRAEG during disease progression and to define stage-specific patterns of molecular evolution relevant to cancer immunotherapy management. METHODS: We analyzed tissue samples from seven patients, including four non-immunotherapy-related gastritis controls and three cancer patients who developed IRAEG following ICI therapy for solid tumors, sampled longitudinally across four disease stages: baseline (G1), acute severe inflammation (G2), early recovery (G3), and complete recovery (G4). Using laser capture microdissection coupled with data-independent acquisition mass spectrometry, we profiled 177 spatially resolved gastric tissue regions. Multiplex immunohistochemistry and immunofluorescence characterized features of the immune microenvironment, while Gene Ontology, KEGG pathway analysis, Gene Set Variation Analysis, and xCell inference enabled functional, metabolic, and immune profiling. Key immune and NET-related findings were further validated by multiplex immunofluorescence in an independent, expanded cohort of IRAEG and non-immunotherapy-related gastritis samples. RESULTS: IRAEG was characterized by widespread HLA molecule activation and enhanced antigen processing, resembling the immune phenotype observed in organ transplant rejection. The acute G2 stage exhibited excessive neutrophil extracellular trap formation, profound metabolic suppression, and collapse of immune homeostasis-features that may inform early intervention strategies to preserve ICI treatment continuity. During early recovery (G3), inflammatory injury transitioned toward repair, marked by activation of fatty acid metabolism and PPAR signaling. Notably, even at complete clinical recovery (G4), more than 1,000 proteins remained differentially expressed, reflecting sustained enhancement of metabolic and immune functions and establishing a distinct molecular "memory" state with implications for ICI rechallenge decisions. CONCLUSIONS: These findings define four molecularly distinct stages of IRAEG progression and recovery. The stage-specific signatures identified here serve as candidate biomarkers for diagnosis, disease staging, and therapeutic response assessment, and may guide clinical decisions regarding irAE management, treatment modification, and safe ICI rechallenge to support continued antitumor therapy.

Humans

Advancing translational exposomics: bridging genome, exposome and personalized medicine.

Understanding the interplay between genetic predisposition and environmental and lifestyle exposures is essential for advancing precision medicine and public health. The exposome, defined as the sum of all environmental exposures an individual encounters throughout their lifetime, complements genomic data by elucidating how external and internal exposure factors influence health outcomes. This treatise highlights the emerging discipline of translational exposomics that integrates exposomics and genomics, offering a comprehensive approach to decipher the complex relationships between environmental and lifestyle exposures, genetic variability, and disease phenotypes. We highlight cutting-edge methodologies, including multi-omics technologies, exposome-wide association studies (EWAS), physiology-based biokinetic modeling, and advanced bioinformatics approaches. These tools enable precise characterization of both the external and the internal exposome, facilitating the identification of biomarkers, exposure-response relationships, and disease prediction and mechanisms. We also consider the importance of addressing socio-economic, demographic, and gender disparities in environmental health research. We emphasize how exposome data can contextualize genomic variation and enhance causal inference, especially in studies of vulnerable populations and complex diseases. By showcasing concrete examples and proposing integrative platforms for translational exposomics, this work underscores the critical need to bridge genomics and exposomics to enable precision prevention, risk stratification, and public health decision-making. This integrative approach offers a new paradigm for understanding health and disease beyond genetics alone.

Humans

Macrolide-resistant Mycoplasma pneumoniae resurgence in Chinese children in 2023: a longitudinal, cross-sectional, genomic epidemiology study.

BACKGROUND: After a prolonged period of low detection rates, Mycoplasma pneumoniae resurged in China, during September to November, 2023, raising global concern. This study aims to gain a better understanding of the genetic mechanisms underlying the 2023 increase in cases and the evolutionary dynamics of the epidemic populations, which has been previously hampered due to limited genomic data of this pathogen. METHODS: We sequenced 685 M pneumoniae isolates, including 248 isolates from 11 Chinese provinces and municipalities in 2023 and 437 isolates from Beijing (2013-22). By analysing these isolates and 436 publicly global sequences, we reconstructed the pathogen's evolutionary history using time-calibrated phylogenies and effective population size inference. We investigated potential genomic variations contributing to the 2023 resurgence through genome-wide association study and conducted phylogeographic analysis of the 2023 isolates across China. FINDINGS: Two macrolide-resistant epidemic clusters (T1-2-EC1 and T2-2-EC2) were responsible for the 2023 resurgence in China. Both clusters, having acquired the 23S ribosomal RNA A2063G mutation conferring macrolide resistance, emerged in approximately 1997 and 2014, respectively, and subsequently outcompeted their predecessor populations. This coincided with China's large-scale adoption of azithromycin for paediatric community-acquired pneumonia around the early 2000s. Aside from macrolide resistance, T1-2-EC1 independently acquired 17 clade-specific mutations and T2-2-EC2 four clade-specific mutations, which could further explain their increased competitiveness. Whole-genome analysis revealed no resurgence-specific mutations in the 2023 isolates. Phylogeographic analysis showed rapid mixing of T1-2-EC1 isolates between different sampled regions within China. INTERPRETATION: Our study provides evidence that the 2023 resurgence in China is a continuation of the pre-COVID epidemic, rather than emergence of novel variants. The high prevalence of macrolide resistance and rapid intranational spread emphasise the urgent need for enhanced global surveillance of this pathogen. FUNDING: National Key Research and Development Program of China, National Natural Science Foundation of China for Key Programs of China Grants, and Beijing High-Level Public Health Technical Talent Project.

Humans

A photochemical method to map ethidium bromide binding sites on DNA: application to a bent DNA fragment.

It is shown that, when irradiated in the visible, ethidium bromide (EB) engages in direct photochemistry with its DNA binding site. At the photochemical end point, an average of one single-strand break is produced per bound EB molecule in a reaction which also bleaches the dye chromophore. Using high-resolution electrophoresis, we have mapped the distribution of EB photocleavage sites on DNA, at one-base resolution. It is argued that because the photocleavage is stoichiometric, the resulting pattern is similar to, if not identical with, the local distribution of EB binding affinity. When interpreted in the context of the extensive thermodynamic and structural data which are available for EB, a binding distribution of that kind can be used to infer details of DNA structure variation within the underlying helix. As a first application of the method, we have used EB to probe the structure of a 265 bp fragment of DNA, which had been described as being bent as the result of a periodic array of oligo(A) segments [Kitchin et al. (1986) J. Biol. Chem. 261, 11302]. The EB mapping data provide evidence that the oligo(A) elements in this fragment assume a local secondary structure which is different than that assumed by isolated ApA nearest neighbors and that the ends of the oligo(A) elements comprise a junctional domain with EB binding properties which differ from those of the oligo(A) element or of random-sequence DNA.

Adenine Nucleotides

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans