Search PubMedSearch

SEARCH · Search PubMed

Results for “parameter estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting analyses are then aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-data analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures, while reducing runtime and memory requirements by orders of magnitude. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical across diverse datasets. We suggest that PSU is a general strategy for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

confidence limits

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework to address this challenge, which reduces runtime and memory requirements by orders of magnitude. In PSU, small subsamples of sites from a concatenated alignment are analyzed, which are expanded by upsampling before inference, and the resulting inferences are aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-alignment analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis of concatenated alignments. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical. We suggest that PSU is a general approach for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

Phylogeny

Comparison of Iodinated Contrast Doses Based on Total Body Weight and Lean Body Weight in Pediatric Patients: Impact on Image Quality and Contrast Exposure.

INTRODUCTION: Iodinated contrast dosing in pediatric computed tomography (CT) traditionally relies on total body weight (TBW), which may result in excessive contrast administration, particularly in patients with higher adiposity. Lean body weight (LBW)-based protocols have shown promise in adults but remain underexplored in children. Therefore, the aim of this study was to compare contrast volume requirements and hepatic enhancement quality among three dosing protocols: LBW-based, TBW-based, and the Control Group (CG), based on the institutional standard for pediatric abdominal CT. METHODS: This prospective study enrolled 66 patients (age 0-16 years) undergoing contrast-enhanced abdominal CT between September 2023 and August 2024. Patients were randomly assigned to receive iodinated contrast (iobitridol 350mg I/mL) dosed by: (1) LBW (0.63 g iodine/kg x LBW, calculated using Peters formula; n = 23), (2) TBW (0.46 g iodine/kg x TBW; n = 20), or (3) institutional control protocol (2 mL/kg x TBW, equivalent to 0.7 g iodine/kg; n = 23). Kruskal-Wallis, ANOVA, Two-way ANOVA, ANCOVA, Scheirer-Ray-Hare, and Cohen's Kappa tests with Likert scale were used. RESULTS: The LBW group received lower median contrast volumes (27 mL; IQR, 10-80 mL) compared to the TBW group (34.5 mL; IQR, 18-78 mL) and the CG group (40 mL; IQR, 13-80 mL), although the differences did not reach statistical significance (P > 0.05). Notably, this reduction did not compromise hepatic enhancement, which remained comparable to the CG (552 ± 139 HU; P = 0.107). CONCLUSION: Lean body weight may be a useful parameter for estimating contrast dose in pediatric abdominal CT, potentially reducing administered volumes without compromising diagnostic image quality. IMPLICATIONS FOR PRACTICE: These results provide early evidence that LBW-based dosing may support more individualized contrast administration in pediatric CT, potentially reducing exposure-related risks.

Humans

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs

Mining Stored-Specimen Studies for Information about Cancer Natural History.

The advent of new multicancer early detection tests and publication of early diagnostic results have generated expectations of clinical benefit from multicancer screening. The clinical benefit of a cancer screening test depends critically on disease natural history, which is typically learned from prospective screening studies. Retrospective studies of stored blood specimens are important in learning about a test's preclinical diagnostic performance but have rarely been used to infer natural history. The extent to which these studies might be harnessed to also learn natural history is discussed in the context of an article in this issue that infers the combined natural history of a range of cancers targeted by a multicancer early detection test using a case-control subsample of specimens from a large cohort study. The critical question concerns the identifiability of key transition rates in multistate models of natural history alongside state-specific sensitivities. The article suggests that these parameters are estimable within a Bayesian framework that leverages prior information about test sensitivity from diagnostic studies. We offer a heuristic discussion of identifiability in this setting and encourage formal study to determine the extent to which models with varying degrees of complexity may be learned from stored-specimen studies. See related article by Dai et al., p. 1535.

Humans

Genomic Insights Into Heterosis: Dominance or Additive × Additive Interaction?

Heterosis was documented in the 18th century, but its biological basis has been debated since. The theoretical framework proposed by Hill, and adapted by Lynch, is based on two central parameters: admixed composition (S), and the heterozygosity (H). Using genomic information, it is now possible to estimate independently the individual realized Si and Hi. In this research, a methodology for estimating the contribution of dominance and additive &#xd7; additive effects to heterosis is proposed. This approach would be especially relevant in cases where there is insufficient phenotypic information available, or an adequate genetic group experimental design, common in humans, wild species and other admixed populations. We also provide theoretical arguments highlighting the enhanced precision of the estimations of heterosis parameters through this method. Furthermore, we exemplify this procedure by analysing data from an experimental F2 pig population, which was initially designed for QTL mapping. Notably, all animals in this population were genotyped (including F1 and parental breeds), but phenotypic information was only available for F2 individuals and included 13 traits related to growth, fat deposition, carcass characteristics and meat quality. Significant additive effects (p&#x2009;<&#x2009;0.05) were detected for longissimus muscle area and carcass temperature, suggesting complementary additive effects for these traits. Significant dominance and additive &#xd7; additive effects were also detected for birth weight and carcass length, respectively (p&#x2009;<&#x2009;0.05), indicating that heterosis for these traits is primarily attributable to dominance and additive &#xd7; additive interactions. These results demonstrate that the proposed methodology can successfully estimate the genetic components underlying heterosis and underscores the utility of this approach in&#xa0;situations where we possess genomic data but limited phenotypic data.

SNP

Integrating genomic additive relationship matrices improves the efficiency in diploid banana breeding.

Partitioning of genetic variance into additive and non-additive components using the pedigree-based best linear unbiased prediction (P-BLUP) model is possible because of the family structure and replicated clones in clonally propagated crops, but this model may overestimate these components. However, the genomic best linear unbiased prediction (G-BLUP) method, which integrates the genetic relationship through molecular marker information reduces the overestimation. Alternatively, a combination of the P-BLUP and G-BLUP, sourcing to create a hybrid matrix that estimates hybrid best linear unbiased prediction (H-BLUP), is proposed. We investigated if integrating molecular information into the clonal model could improve the partitioning of the variance components leading to more accurate estimates of genetic parameters and prediction accuracy of breeding values of 14 key traits in diploid banana. In this study, we used clones of 14 full-sib families from a factorial mating design of four female and five diploid male banana (Musa acuminata) parents, generated at the International Institute of Tropical Agriculture in Arusha. The genomic-based relationship matrices were constructed using a set of 2792 filtered single-nucleotide polymorphism markers. Additive variance and heritability derived from G-BLUP and H-BLUP models reduced bias compared to the P-BLUP model. The H-BLUP estimated the highest prediction accuracies for yield-related and cycling traits, while the P-BLUP model had the highest prediction accuracy estimates for agronomic traits. The use of marker-based models enhances the accuracy of predicting breeding values, contributing to accurate estimates of genetic gain while paving a way for further genomic exploration in diploid banana breeding programs.

Journal Article

A doubly robust framework for addressing outcome-dependent selection bias in multi-cohort EHR studies.

Selection bias can hinder accurate estimation of association parameters in binary disease risk models using non-probability samples like electronic health records (EHRs). The issue is compounded when participants are recruited from multiple clinics/centers with varying selection mechanisms that may depend on the disease/outcome of interest. Traditional inverse-probability-weighted (IPW) methods, based on constructed parametric selection models, often struggle with misspecifications when selection mechanisms vary across cohorts. This paper introduces a new Joint Augmented Inverse Probability Weighted (JAIPW) method, which integrates individual-level data from multiple cohorts collected under potentially outcome-dependent selection mechanisms, with data from an external probability sample. JAIPW offers double robustness by incorporating a flexible auxiliary score model to address potential misspecifications in the selection models. We outline the asymptotic properties of the JAIPW estimator, and our simulations reveal that JAIPW achieves up to 6 times lower relative bias and 5 times lower root mean square error (RMSE) compared to the best performing joint IPW methods under scenarios with misspecified selection models. Applying JAIPW to the Michigan Genomics Initiative (MGI), a multi-clinic EHR-linked biobank, combined with external national probability samples, resulted in cancer-sex association estimates closely aligned with national benchmark estimates. We also analyzed the association between cancer and polygenic risk scores (PRS) in MGI to illustrate a situation where the exposure variable is not measured in the external probability sample.

Selection Bias

Chronotype and cellular circadian rhythms predict the clinical response to lithium maintenance treatment in patients with bipolar disorder.

Bipolar disorder (BD) is a serious mood disorder associated with circadian rhythm abnormalities. Risk for BD is genetically encoded and overlaps with systems that maintain circadian rhythms. Lithium is an effective mood stabilizer treatment for BD, but only a minority of patients fully respond to monotherapy. Presently, we hypothesized that lithium-responsive BD patients (Li-R) would show characteristic differences in chronotype and cellular circadian rhythms compared to lithium non-responders (Li-NR). Selecting patients from a prospective, multi-center, clinical trial of lithium monotherapy, we examined morning vs. evening preference (chronotype) as a dimension of circadian rhythm function in 193 Li-R and Li-NR BD patients. From a subset of 59 patient&#xa0;donors, we measured circadian rhythms in skin&#xa0;fibroblasts longitudinally over 5 days using a bioluminescent reporter (Per2-luc). We then estimated circadian rhythm parameters (amplitude, period, phase) and the pharmacological effects of lithium on rhythms in cells from Li-R and Li-NR donors. Compared to Li-NRs, Li-Rs showed a difference in chronotype, with higher levels of morningness. Evening chronotype was associated with increased mood symptoms at baseline, including depression, mania, and insomnia. Cells from Li-Rs were more likely to exhibit a short circadian period, a linear relationship between period and phase, and period shortening effects of lithium. Common genetic variation in the IP3 signaling pathway may account for some of the individual differences in the effects of lithium on cellular rhythms. We conclude that circadian rhythms may influence response to lithium in maintenance treatment of BD.

Adult

GeomeTRe: accurate calculation of geometrical descriptors of tandem repeat proteins.

MOTIVATION: Structured tandem repeat proteins (STRPs) are characterized by preserved structural motifs arranged in a modular way. The structural and functional diversity of STRPs makes them particularly important for studying evolution and novel structure-function relationships, and ultimately for designing new synthetic proteins with specific functions. One crucial aspect of their classification is the estimation of geometrical parameters, which can provide better insight into their properties and the relationship between the spatial arrangement of repeated units and protein function. Calculating geometric descriptors for STRPs is challenging because naturally occurring repeats are not "perfect" and often contain insertions and deletions. Existing tools for predicting structural symmetry work well on simple cases but often fail for most natural proteins. RESULTS: Here, we present GeomeTRe, an algorithm that calculates geometrical descriptors such as curvature (yaw), twist (roll), and pitch for a protein structure with known repeat unit positions. The algorithm simulates the movement of consecutive units, identifies rotational axes, and calculates the corresponding Tait-Bryan angles. GeomeTRe's parameters can enhance STRP annotation and classification by identifying variations in geometric arrangements among different functional groups. The package is fast and suitable for processing large protein structure datasets when repeat region information (e.g. from RepeatsDB) is available. AVAILABILITY AND IMPLEMENTATION: GeomeTRe is available as a Python package; source code and documentation can be found at https://github.com/BioComputingUP/GeomeTRe.

Algorithms

Impact of genomic selection for disease resistance on the spread of infection in a simulated aquaculture population.

BACKGROUND: In aquaculture, selection for disease resistance is typically based on mortality records from challenge tests performed on relatives of selection candidates. However, commercial success depends on limiting disease transmission, particularly the incidence and severity of outbreaks. It remains unclear whether selecting for lower mortality also reduces disease transmission. Both these outcomes are influenced by three underlying epidemiological host traits: susceptibility, infectivity, and infection-induced mortality. This simulation study evaluated the impact of genomic selection against mortality on disease transmission in a salmon population exposed to a pathogen with a fast transmission rate. METHODS: Mortality was assumed to be recorded on sibs of selection candidate, either as binary dead/alive status or as time to death during cohabitation/bath challenge tests. Phenotypes were simulated using a stochastic compartmental Susceptible-Infected-Removed epidemiological model, with genetic variation for the three underlying traits. Scenarios were explored by varying the genetic correlations between the three underlying traits. Challenge test designs varied in the number of groups, group sizes, and family distribution across groups. For comparison, a reference scenario with direct selection on the underlying traits was included. Genomic selection was applied over 10 discrete generations, and its impact on disease transmission was assessed using the basic reproductive ratio (R0). RESULTS: When selection was based on dead/alive status, R0 was highly sensitive to both the genetic correlations between the three underlying traits and the challenge test design. In contrast, selection based on time to death consistently reduced R0 across all scenarios (often to below 1 within four generations), regardless of trait correlations or test design. Selection on time to death primarily produced fish with reduced susceptibility to infection, while selection on dead/alive status produced fish with increased resistance and endurance to infection, delaying onset of infection or death without necessarily limiting transmission. Direct selection on the underlying epidemiological traits was the most efficient approach to reduce both mortality and transmission. CONCLUSIONS: Genomic selection for disease resistance, when measured as time to death in cohabitation or bath challenge tests conducted until mortality naturally levels off, reduces both mortality and disease spread. Breeding programs may benefit from challenge test designs that enable estimation of genetic parameters for the underlying traits affecting disease transmission and survival.

Animals

Bayesian identification of differentially expressed isoforms using a novel joint model of RNA-seq data.

We develop a Bayesian approach, BayesIso, to identify differentially expressed isoforms from RNA-seq data. The approach features a novel joint model of the sample variability and the deferential state of isoforms. Specifically, the within-sample variability and the between-sample variability of each isoform are modeled by a Poisson-Lognormal model and a Gamma-Gamma model, respectively. Using a Bayesian framework, the differential state of each isoform and the model parameters are jointly estimated by a Markov Chain Monte Carlo (MCMC) method. Extensive studies using simulation and real data demonstrate that BayesIso can effectively detect isoforms of less differentially expressed and differential transcripts for genes with multiple isoforms. We applied the approach to breast cancer RNA-seq data and uncovered a unique set of isoforms that form key pathways associated with breast cancer recurrence. First, PI3K/AKT/mTOR signaling and PTEN signaling pathways are identified as being involved in breast cancer development. Further integrated with protein-protein interaction data, pathways of Jak-STAT, mTOR, MAPK and Wnt signaling are revealed in association with breast cancer recurrence. Finally, several pathways are activated in the early recurrence of breast cancer. In tumors that occur early, members of pathways of cellular metabolism and cell cycle (such as CD36 and TOP2A) are upregulated, while immune response genes such as NFATC1 are downregulated.

Humans

Inference of elevated mutation rates and variant effects using 700k exomes.

Genomic sequencing is now widely accessible for genetic diagnostics and is emerging as a component of newborn screening. This technological development generates the need to characterize incoming mutations, create comprehensive datasets of genes causing rare Mendelian disorders, and identify pathogenic variants. Large-scale exome sequencing datasets such as Genome Aggregation Database (gnomAD) have been assembled to help address these challenges. The recent release of gnomAD (v4; n = 730,947) uncovers millions of rare coding variants, many of which have arisen more than once by independent recurrent mutations in the rapidly growing recent human population. Here, we use newly developed theoretical understanding of sampling properties of rare variants to estimate key population genetics parameters of practical importance to human genetics such as demography history, mutation rate, and selection. Solely relying on population data, our method Population Inferred Estimates of Selection (PIES) identifies novel genes with loss-of-function mutational hotspots likely due to selection in spermatogonia. PIES efficiently estimates selection coefficients for heterozygous loss-of-function variants. Combining population genetics inference with variant effect predictors, PIES predicts pathogenic missense mutations and improves variant prioritization for genetic diagnostics and newborn screening.

Journal Article

Striping artifact removal in VisiumHD data through nuclear counts modeling.

MOTIVATION: 10x Genomics VisiumHD enables spatial transcriptomics at 2&#x2009;&#xb5;m &#xd7; 2&#x2009;&#xb5;m resolution but exhibits slide-specific, non-periodic striping artifacts due to lane-width variability. These multiplicative row/column effects distort bin total counts and can bias downstream analyses. The state-of-the-art destriping approach is the normalization procedure used as a preprocessing step in bin2cell; it applies sequential high-quantile row- then column-wise normalization, which is asymmetric and can introduce edge effects/macro-stripes and distortions of large-scale total-count structure. RESULTS: We propose a statistical destriping approach that leverages nuclei segmentation from the co-registered H&E image. Assuming transcript abundance is constant within each nucleus, we model bin counts with a negative binomial distribution whose mean is a product of a nucleus-specific concentration and row- and column-specific stripe-factors reflecting lane-width variation. We fit all parameters in a generalized linear modeling framework with cross-validated regularization on stripe-factors and iterative dispersion estimation, and use the fitted parameters to correct the observed counts into a destriped image. On synthetic data with known ground truth, our method improves stripe-factor estimation accuracy and reduces error in corrected counts relative to bin2cell and bin2cell-derived baselines. Across four public VisiumHD slides, it consistently lowers striping intensity while substantially better preserving biological signal present in the large-scale global count structure and avoiding the artifacts introduced by other methods. AVAILABILITY AND IMPLEMENTATION: All source code and links to publicly available data used for this study are available at https://github.com/paolamalsot/destriping-GLM.

Artifacts

Dissecting fluctuating selection: A unified population and quantitative genetics framework.

One of the longstanding debates in evolutionary biology is the effect of fluctuating selection on genetic changes in populations. However, the extent to which these periodic forces influence organisms at both genomic and phenotypic levels remains unclear. Despite the compelling evidence of fluctuating selection from recent studies, there is a disconnect between empirical and theoretical findings concerning the underlying mechanisms due to the limited evidence regarding the scale and processes that generate genome-wide oscillations. This study aims to elucidate how both genetic factors (e.g. heritability, number of causative loci) and ecological factors (e.g. season length, the difference in the phenotypic optima between seasons, population size dynamics) drive fluctuating selection and to identify the parameters that produce consistent oscillatory patterns. We developed a modeling framework integrating quantitative and population genetics to simulate a population under various selection regimes. We applied spectral analysis to detect periodicity, indicating cyclical selective environments. Our simulations highlight the conditions sustaining oscillations in allele frequencies over time. Spectral analysis successfully identifies the periodic patterns from allele frequency trajectories, even under highly complex selection regimes. Not only does our study clarify the conditions that yield oscillatory behaviors, but these parameters can also potentially be estimated in natural populations, providing a possibility of empirically testing these models.

Fluctuating selection

vcfsim: flexible simulation of all-sites VCFs with missing data.

BACKGROUND |: VCFs are the most widely used data format for encoding genetic variation. By design, standard VCFs do not include data from sites where all individuals are homozygous for the reference allele ("invariant sites") and thus do not differentiate these from sites where data are completely missing. However, missing data are a key feature of biological datasets across all domains of genomics, and many recent studies have shown that missing data can introduce a variety of statistical biases in the estimation of key population genetic parameters. A solution to this limitation is to include invariant sites in a standard VCF, creating an "all-sites VCF", exposing missing and invariant sites explicitly. One hurdle to the wider adoption of all-sites VCFs is a reliable parameterized simulation framework for generating biologically realistic all-sites VCFs. RESULTS |: Here, we introduce an open-source command line tool, vcfsim, that interfaces with the popular coalescent simulation platform msprime and provides convenience functions for simulating all-sites VCFs with variable levels of ploidy and missing data. We show that the post-processed VCFs generated using vcfsim align precisely with population genetic expectations (i.e. are statistically identical to raw msprime output), accurately introduce missing data, and permit the simulation of data with varying ploidy levels, including the simulation of intraindividual ploidy variation (e.g. heterogametic sex chromosomes) and population structures. CONCLUSIONS |: Our results vcfsim is a useful and easy-to-use tool for the benchmarking of new software tools, performing population genetic inference, training of machine learning models, and the exploration of the effects of missing data in genomics data sets.

Benchmarking

Dissecting fluctuating selection: A unified population and quantitative genetics framework.

One of the longstanding debates in evolutionary biology is the effect of fluctuating selection on genetic changes in populations. However, the extent to which these periodic forces influence organisms at both genomic and phenotypic levels remains unclear. Despite the compelling evidence of fluctuating selection from recent studies, there is a disconnect between empirical and theoretical findings concerning the underlying mechanisms due to the limited evidence regarding the scale and processes that generate genome-wide oscillations. This study aims to elucidate how both genetic factors (e.g. heritability, number of causative loci) and ecological factors (e.g. season length, the difference in the phenotypic optima between seasons, population size dynamics) drive fluctuating selection and to identify the parameters that produce consistent oscillatory patterns. We developed a modeling framework integrating quantitative and population genetics to simulate a population under various selection regimes. We applied spectral analysis to detect periodicity, indicating cyclical selective environments. Our simulations highlight the conditions sustaining oscillations in allele frequencies over time. Spectral analysis successfully identifies the periodic patterns from allele frequency trajectories, even under highly complex selection regimes. Not only does our study clarify the conditions that yield oscillatory behaviors, but these parameters can also potentially be estimated in natural populations, providing a possibility of empirically testing these models.

Fluctuating selection

Estimating Re and overdispersion in secondary cases from the size of identical sequence clusters of SARS-CoV-2.

The wealth of genomic data that was generated during the COVID-19 pandemic provides an exceptional opportunity to obtain information on the transmission of SARS-CoV-2. Specifically, there is great interest to better understand how the effective reproduction number [Formula: see text] and the overdispersion of secondary cases, which can be quantified by the negative binomial dispersion parameter k, changed over time and across regions and viral variants. The aim of our study was to develop a Bayesian framework to infer [Formula: see text] and k from viral sequence data. First, we developed a mathematical model for the distribution of the size of identical sequence clusters, in which we integrated viral transmission, the mutation rate of the virus, and incomplete case-detection. Second, we implemented this model within a Bayesian inference framework, allowing the estimation of [Formula: see text] and k from genomic data only. We validated this model in a simulation study. Third, we identified clusters of identical sequences in all SARS-CoV-2 sequences in 2021 from Switzerland, Denmark, and Germany that were available on GISAID. We obtained monthly estimates of the posterior distribution of [Formula: see text] and k, with the resulting [Formula: see text] estimates slightly lower than estimates obtained by other methods, and k comparable with previous results. We found comparatively higher estimates of k in Denmark which suggests less opportunities for superspreading and more controlled transmission compared to the other countries in 2021. Our model included an estimation of the case detection and sampling probability, but the estimates obtained had large uncertainty, reflecting the difficulty of estimating these parameters simultaneously. Our study presents a novel method to infer information on the transmission of infectious diseases and its heterogeneity using genomic data. With increasing availability of sequences of pathogens in the future, we expect that our method has the potential to provide new insights into the transmission and the overdispersion in secondary cases of other pathogens.

COVID-19