Search PubMedSearch

SEARCH · Search PubMed

Results for “Bayesian estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Bayesian estimation of allele-specific expression in the presence of phasing uncertainty.

MOTIVATION: Allele-specific expression (ASE) analyses aim to detect imbalanced expression of maternal versus paternal copies of an autosomal gene. Such allelic imbalance can result from a variety of cis-acting causes, including disruptive mutations within one copy of a gene that impact the stability of transcripts, as well as regulatory variants outside the gene that impact transcription initiation. Current methods for ASE estimation suffer from a number of shortcomings, such as relying on only one variant within a gene, assuming perfect phasing information across multiple variants within a gene, or failing to account for alignment biases and possible genotyping errors. RESULTS: We developed BEASTIE, a Bayesian hierarchical model designed for precise ASE quantification at the gene level, based on given genotypes and RNA-Seq data. BEASTIE addresses the complexities of allelic mapping bias, genotyping error, and phasing errors by incorporating empirical phasing error rates derived from Genome-in-a-Bottle individual NA12878. BEASTIE surpasses existing methods in accuracy, especially in scenarios with high phasing errors. This improvement is critical for identifying rare genetic variants often obscured by such errors. Through rigorous validation on simulated data and application to real data from the 1000 Genomes Project, we establish the robustness of BEASTIE. These findings underscore the value of BEASTIE in revealing patterns of ASE across gene sets and pathways. AVAILABILITY AND IMPLEMENTATION: The software is freely available from Github (https://github.com/x811zou/BEASTIE); and Zendo (DOI: 10.5281/zenodo.15062124).

Bayes Theorem

NExON-Bayes: a Bayesian approach to network estimation informed by ordinal covariates.

MOTIVATION: In heterogeneous disease settings, accounting for intrinsic sample variability is crucial for obtaining reliable and interpretable omic network estimates. However, most graphical model analyses of biomedical data assume homogeneous conditional dependence structures, potentially leading to misleading conclusions. To address this, we propose a joint Gaussian graphical model that leverages sample-level ordinal covariates (e.g. disease stage) to account for heterogeneity and improve the estimation of partial correlation structures. RESULTS: Our modelling framework, called NExON-Bayes, extends the graphical spike-and-slab framework to account for ordinal covariates, jointly estimating their relevance to the graph structure and leveraging them to improve the accuracy of network estimation. To scale to high-dimensional omic settings, we develop an efficient variational inference algorithm tailored to our model. Through simulations, we demonstrate that our method outperforms the vanilla graphical spike-and-slab (with no covariate information), as well as other state-of-the-art network approaches which exploit covariate information. Applying our method to reverse phase protein array data from patients diagnosed with stage I, II or III breast carcinoma, we estimate the behaviour of proteomic networks as cancer progresses. Our model provides insights not only through inspection of the estimated proteomic networks, but also of the estimated ordinal covariate dependencies of key groups of proteins within those networks, offering a comprehensive understanding of how biological pathways shift across disease stages. AVAILABILITY AND IMPLEMENTATION: A user-friendly R package for NExON-Bayes with tutorials is available on Github at github.com/jf687/NExON, and archived at https://doi.org/10.5281/zenodo.20312938. The source of the dataset used is cited in the relevant section.

Bayes Theorem

A novel second-generation HIV-1 circulating recombinant form (CRF183_0107) identified among men who have sex with men in China.

OBJECTIVE: This study aimed to report a novel HIV-1 circulating recombinant form (CRF) identified among men who have sex with men (MSM) in China. DESIGN: Viral sequences were isolated from MSM patients, and the recombination and evolutionary histories of this CRF were elucidated through phylogenetic and Bayesian analyses. METHODS: Near full-length genomes (NFLGs) and partial genome sequences were amplified from RNA extracted from plasma samples of three HIV-1 seropositive MSM in Heilongjiang Province, China. Phylogenetic analysis was conducted using FastTree v2.1.9, and recombination analysis was performed using Simplot v3.5.1. The emergence time of the novel CRF was estimated by Bayesian evolutionary analysis using BEAST v1.10.4. RESULTS: Two NFLGs and two partial genome segments were successfully obtained from three MSM participants. This novel CRF was characterized by 12 mosaic gene segments, comprising 6 segments from the CRF01_AE cluster 4 and 6 segments from the CRF07_BC cluster N, and thus was designated as CRF183_0107. The estimated time of origin for the CRF01_AE and CRF07_BC components within CRF183_0107 were approximately 2009.2 and 2012.5, respectively. CONCLUSION: A novel second-generation HIV-1 recombinant, named CRF183_0107, was identified within the MSM population in China. This CRF exemplified the recombination events occurring between the CRF01_AE cluster 4 and the CRF07_BC cluster N during 2009-2012.

Humans

Bayesian Modeling of Cancer Outcomes Using Genetic Variables Assisted by Pathological Imaging Data.

With the increasing maturity of genetic profiling, an essential and routine task in cancer research is to model disease outcomes/phenotypes using genetic variables. Many methods have been successfully developed. However, oftentimes, empirical performance is unsatisfactory because of a "lack of information." In cancer research and clinical practice, a source of information that is broadly available and highly cost-effective comes from pathological images, which are routinely collected for definitive diagnosis and staging. In this article, we consider a Bayesian approach for selecting relevant genetic variables and modeling their relationships with a cancer outcome/phenotype. We propose borrowing information from (manually curated, low-dimensional) pathological imaging features via reinforcing the same selection results for the cancer outcome and imaging features. We further develop a weighting strategy to accommodate the scenario where information borrowing may not be equally effective for all subjects. Computation is carefully examined. Simulations demonstrate competitive performance of the proposed approach. We analyze TCGA (The Cancer Genome Atlas) LUAD (lung adenocarcinoma) data, with overall survival and gene expressions being the outcome and genetic variables, respectively. Findings different from the alternatives and with sound properties are made.

Humans

Efficacy and safety of cannabinoid-based interventions for behavioral and cognitive symptoms in dementia: systematic review and meta-analysis.

BACKGROUND: Behavioral and cognitive symptoms are frequent in Alzheimer's disease and dementia, and available pharmacological options offer limited benefit. Cannabinoid-based therapies have been proposed as alternatives, but evidence remains inconclusive. METHODS: We systematically searched PubMed, Embase, Web of Science, and the Cochrane Library through November 2025 for randomized controlled trials evaluating cannabinoids in Alzheimer's disease or dementia. Primary outcomes were agitation measured by the Cohen-Mansfield Agitation Inventory (CMAI) and neuropsychiatric symptoms assessed by the Neuropsychiatric Inventory-Nursing Home version (NPI-NH). Secondary outcomes included cognition using the Mini-Mental State Examination (MMSE) and adverse events. Standardized Mean Differences (SMDs) and Risk Ratios (RRs) were synthesized using random-effects (REML) and Bayesian random-effects models. Risk of bias was evaluated with RoB 2, and certainty of evidence with GRADE. RESULTS: Nine trials (334 participants) met inclusion criteria. Cannabinoids did not improve CMAI (SMD -0.58, 95% CI -1.71 to 0.55; I2 = 84%), NPI-NH total (SMD -0.02, 95% CI -1.00 to 0.96; I2 = 67%), NPI-NH agitation (SMD -0.44, 95% CI -1.45 to 0.57; I2 = 48%), or MMSE (SMD 0.86, 95% CI -16.33 to 18.06; I2 = 96%). Bayesian posterior estimates were close to zero, supporting the absence of effect. Leave-one-out analyses reduced heterogeneity only after excluding influential trials but did not alter results. Certainty of evidence was moderate for behavioral outcomes and low for cognition. Overall adverse events were similar to placebo, while somnolence was more frequent with cannabinoids (RR 2.03, 95% CI 1.29-3.20). CONCLUSIONS: Cannabinoid-based therapies do not improve agitation, neuropsychiatric symptoms, or cognition in Alzheimer's disease and increase somnolence.

Humans

Plasmacytoid dendritic cell-mediated L-glutamate catabolism links gut microbiota to male infertility.

Emerging evidence suggests that gut microbiota composition influences male reproductive health; however, the immunometabolic mechanisms underlying this association remain insufficiently characterized. We investigated whether specific immune cell-mediated metabolic pathways, particularly plasmacytoid dendritic cell (pDC)-driven L-glutamate catabolism via the hydroxyglutarate pathway, contribute to the causal link between gut microbiota and male infertility. We conducted a 2-sample, 2-step Mendelian randomization (MR) analysis using inverse-variance weighting as the primary estimator and Bayesian weighted MR for robustness. Exposure data comprised 412 gut microbial taxa/metabolic pathways and 731 immune cell phenotypes from large European-ancestry genome-wide association studies. Male infertility genome-wide association studies data (1429 cases; 128,710 controls) were obtained from FinnGen R10. Only exposure-mediator-outcome pairs meeting stringent pleiotropy, heterogeneity, and reverse-causality criteria were retained for mediation analysis. Nine microbial taxa/metabolic pathways and 18 immune traits exhibited putative causal associations with male infertility. The L-glutamate degradation V pathway via hydroxyglutarate was linked to reduced infertility risk (inverse-variance weighting odds ratio [OR] = 0.68; 95% confidence interval, 0.52-0.89; P = .005). Two-step MR suggested that forward scatter area on pDCs may mediate this association, although the mediation effect was imprecise (effect = 0.0277; 95% confidence interval, -0.0348 to 0.0903). This study provides suggestive genetic evidence that pDC-mediated glutamate catabolism may connect gut microbial metabolic activity to male infertility. These findings highlight immunometabolic pathways as testable targets for mechanistic validation and microbiota-directed interventions.

Male

Identification of a novel HIV-1 circulating recombinant form (CRF209_cpx) and its descendant unique recombinant form (URF) CRF209_cpx/B among MSM in Guangdong, southern China.

BACKGROUND: The epidemic of human immunodeficiency virus type 1 (HIV-1) continues to pose a significant global health challenge, with increasing genetic diversity. The co-circulation of multiple subtypes among the local population facilitates the emergence of unique or circulating recombinant forms (URFs or CRFs). In China, the predominant strains include CRF07_BC, CRF01_AE, CRF55_01B, and subtype B. This study characterizes a novel CRF209_cpx and its descendant recombinant CRF209_cpx/B among men who have sex with men (MSM) in Guangdong, southern China. METHODS: Individuals infected with URFs with similar genetic characteristics were recruited during routine surveillance of pretreatment drug resistance. Near full-length genomes (NFLGs) were amplified with two overlapping fragments using a serial dilution nested PCR approach after reverse transcription. We used SimPlot and IQ-TREE softwares to conduct recombination analyses and phylogenetic inferences. Time-scaled maximum clade credibility (MCC) phylogenetic trees were reconstructed using BEAST software to estimate evolutionary origins. Genotypic drug resistance mutations were interpreted via the Stanford HIV Database, and coreceptor usage was predicted using geno2pheno coreceptor 2.5 and the HIVcoPRED tool. RESULTS: Four NFLG sequences were obtained and identified as a novel CRF209_cpx, generated by recombination among CRF01_AE, CRF07_BC and subtype B. Phylogenetic analyses revealed that all the parental segments clustered with lineages prevalent among MSM in China. Bayesian evolutionary analysis estimated that the most recent common ancestor (tMRCA) of CRF209_cpx to have evolved between 2011 and 2013. The fifth strain was identified as a URF recombined from nascent CRF209_cpx and B. No transmitted drug resistance mutation was detected in these five sequences. The four CRF209_cpx sequences primarily utilized the CXCR4 coreceptor, while the URF exhibited R5/X4 dual tropism. CONCLUSIONS: The emergence of the complex CRF209_cpx and novel URF of CRF209_cpx/B highlights the active HIV-1 epidemic within the MSM population in Guangdong, underscoring the necessity for enhanced molecular surveillance and precise public health intervention in this key population.

HIV-1

Inference of Gene Flow between Species from Genomic Data When the Mode, Direction, and Lineages are Misspecified.

Thanks to genomic data, interspecific gene flow is increasingly recognized as a major evolutionary force that shapes biodiversity. Two models have been developed in the multispecies coalescent (MSC) framework to infer gene flow from genomic data, assuming either constant-rate continuous migration (MSC-M) or discrete introgression/hybridization (MSC-I). The extreme simplicity of these models raises concerns about their usefulness as they represent misspecified models when applied to real data. Here, we study inference of gene flow under the MSC-M model, considering mis-assignment of gene flow onto incorrect parental or daughter lineages, misspecification of the direction of gene flow, and misspecification of the mode of gene flow. Mis-assignment of gene flow to an incorrect lineage causes large biases in the estimated rates. The Bayesian test has high power for inferring both recent and ancient gene flow, between either sister lineages or nonsister lineages, although misspecification of the direction of gene flow may make it hard to distinguish early divergence with gene flow from recent complete isolation. Misspecification of the mode of gene flow (MSC-I versus MSC-M) has small local effects, and gene flow is detected with high power despite the misspecification. We analyze a genomic dataset from the purple cone spruce (Picea spp., Pinaceae), which putatively arose through homoploid hybrid speciation, to demonstrate practical implications of our theoretical analyses. Overall, we find that the extremely idealized models of gene flow (in particular the discrete MSC-I model) are very effective for extracting information about species divergence and gene flow from genomic data.

Gene Flow

Mining Stored-Specimen Studies for Information about Cancer Natural History.

The advent of new multicancer early detection tests and publication of early diagnostic results have generated expectations of clinical benefit from multicancer screening. The clinical benefit of a cancer screening test depends critically on disease natural history, which is typically learned from prospective screening studies. Retrospective studies of stored blood specimens are important in learning about a test's preclinical diagnostic performance but have rarely been used to infer natural history. The extent to which these studies might be harnessed to also learn natural history is discussed in the context of an article in this issue that infers the combined natural history of a range of cancers targeted by a multicancer early detection test using a case-control subsample of specimens from a large cohort study. The critical question concerns the identifiability of key transition rates in multistate models of natural history alongside state-specific sensitivities. The article suggests that these parameters are estimable within a Bayesian framework that leverages prior information about test sensitivity from diagnostic studies. We offer a heuristic discussion of identifiability in this setting and encourage formal study to determine the extent to which models with varying degrees of complexity may be learned from stored-specimen studies. See related article by Dai et al., p. 1535.

Humans

Spatially Contextualized Integrative Genomics Highlights Neuronal and Glial Regulatory Programs in Low Back Pain.

PURPOSE: Low back pain (LBP) is a heterogeneous pain condition with a measurable genetic contribution, but the genes, brain cell types, and spatial tissue contexts through which inherited risk is expressed remain unclear. We aimed to define cell-type-specific and spatially contextualized genetic mechanisms underlying LBP. METHODS: FinnGen R12 LBP GWAS summary statistics (42,521 cases and 353,224 controls) were integrated with brain single-nuclei eQTL data across eight major brain cell classes. We evaluated genome-wide polygenic signal using LDSC, prioritized genes using MAGMA and PoPS, and performed brain cell-type-specific eQTL-anchored Mendelian randomization, primarily based on single-instrument Wald ratio estimates, followed by Bayesian colocalization. Spatial genetic mapping was conducted using gsMap in an E16.5 murine embryonic atlas and two adult human lumbar spinal cord Visium sections. Selected candidates were assessed by RT-qPCR in neuronal-like and astroglial-like inflammatory cell models. RESULTS: LDSC supported interpretable polygenic signal for LBP. MAGMA and PoPS showed partial gene-level convergence, with TCF4 and TMEFF2 supported by both approaches. Across 1641 tested gene-cell type exposures, significant eQTL-anchored MR associations were concentrated in excitatory neurons, oligodendrocytes, inhibitory neurons, and astrocytes. Integrated eQTL-anchored MR, colocalization, and gene-prioritization evidence highlighted CLEC18A, QPRT, and GMPPB as higher-priority non-MHC candidates with moderate, but not strong, colocalization support. gsMap localized LBP-associated enrichment to neuroaxis-related embryonic regions, including brain, spinal cord, sympathetic nerve, and dorsal root ganglion, and to neuronal-like niches in adult lumbar spinal cord. RT-qPCR showed model-dependent expression changes, with QPRT and LGI4 preferentially responsive in neuronal-like SH-SY5Y cells and GMPPB and DPYSL5 responsive in astroglial-like U251 cells. CONCLUSION: These findings support neuronal and glial regulatory programs as plausible contributors to LBP genetic susceptibility and highlight CLEC18A, QPRT, and GMPPB as higher-priority non-MHC candidates with moderate colocalization support. The results provide a spatially contextualized framework for candidate prioritization in LBP, while emphasizing the need for larger cell-type-specific eQTL resources and functional validation before therapeutic or mechanistic conclusions can be drawn.

Mendelian randomization

Inferring the sensitivity of wastewater metagenomic sequencing for early detection of viruses: a statistical modelling study.

BACKGROUND: Metagenomic sequencing of wastewater (W-MGS) can in principle detect any known or novel pathogen in a population. We aimed to quantify the sensitivity and cost of W-MGS for viral pathogen detection by jointly analysing W-MGS and epidemiological data for a range of human-infecting viruses. METHODS: In this statistical modelling study, we analysed sequencing data from four studies of untargeted W-MGS to estimate the relative abundance of 11 human-infecting viruses. Corresponding prevalence and incidence estimates were obtained or calculated from academic and public health reports. We combined these estimates using a hierarchical Bayesian model to predict relative abundance at set prevalence or incidence values, allowing comparison across studies and viruses. These predictions were then used to estimate the sequencing depth and concomitant cost required for pathogen detection using W-MGS with or without use of a hybridisation capture enrichment panel. FINDINGS: After controlling for variation in local infection rates, relative abundance varied by orders of magnitude across studies for a given virus. For instance, a local SARS-CoV-2 weekly incidence of 1% corresponded to a predicted SARS-CoV-2 relative abundance ranging from 3·8 × 10-10 to 2·4 × 10-7 across studies, translating to orders-of-magnitude variation in the cost of operating a system able to detect a SARS-CoV-2-like pathogen at a given sensitivity. Use of a respiratory virus enrichment panel in two studies greatly increased predicted relative abundance of SARS-CoV-2, lowering yearly costs by 27-fold (from US$7·87 million to $287 000) and 29-fold (from $1·98 million to $69 100) for a system able to detect a SARS-CoV-2-like pathogen before reaching 0·01% cumulative incidence. INTERPRETATION: The large variation in viral relative abundance after controlling for epidemiological factors indicates that other sources of inter-study variation, such as differences in sewershed hydrology and laboratory protocols, have a substantial impact on the sensitivity and cost of W-MGS. Well chosen hybridisation capture panels can greatly increase sensitivity and reduce cost for viruses in the panel, but might reduce sensitivity to unknown or unexpected pathogens. FUNDING: The Wellcome Trust, Open Philanthropy, and Musk Foundation.

Humans

Parting ways: Pan-Homo divergence revisited.

The timing of divergence between hominins and the bonobo-chimpanzee clade has been at the core of palaeoanthropological debate for over a century. The earliest molecular studies indicated divergence times ranging from 5 Ma to as recently as 1.3 Ma. This study critically reviews the trends of time estimates published between 1967 and 2023, and analyses how these are supported or rejected by the current molecular and fossil records. We compiled 202 divergence estimates and defined three distinct thresholds based on fossil evidence at 4.4 Ma (Australopithecus anamensis and Ardipithecus ramidus), 6.2 Ma (Orrorin tugenensis and Ardipithecus kadabba), and 7.2 Ma (Sahelanthropus tchadensis). We then used these thresholds to filter out molecular estimates that are too young to fit the fossil record. Overall, the data suggests a divergence event within the late Miocene, with each threshold pushing it further back, 8.63-6.38, 10.33-7.81, and 10.95-8.81 Ma, respectively. We use a quadratic regression to demonstrate that estimates have been slowly shifting from ~ 6 Ma to ~ 8.5 Ma over the past 56 years. A Bayesian meta-analysis of genomic estimates filtered by our most consensual threshold (i.e., assuming Australopithecus belongs to Hominini) indicates that the split must have occurred early in the late Miocene, most likely before 7 Ma (~ 99.5% posterior probability) with a pooled effect of 8.69-7.28 Ma. We conclude that, despite an initial bias towards younger estimates, the molecular timing for the last common ancestor (LCA) of Pan-Homo has been progressively approaching the intervals suggested by the current fossil record.

Animals

Contrasting Patterns of Connectivity Between Populations of Euphotic and Mesophotic Hydroids in Reunion Island Support the Deep Reef Refuge Hypothesis.

In the context of coral reef decline, mesophotic coral ecosystems (MCEs, 30-150 m) offer hope for the recovery of degraded euphotic reefs. The Deep Reef Refuge Hypothesis (DRRH) postulates the potential of mesophotic reefs to reseed euphotic reefs. This hypothesis needs to be further tested by estimating connectivity along the depth gradient. Mesophotic data are lacking worldwide, particularly in the southwestern Indian Ocean (SWIO). Here, using a total of 2218 samples collected at depths ranging from 10 to 103 m, we estimated the connectivity of 7 hydroid species sampled at euphotic, upper, and lower mesophotic depths around Reunion Island using a multi-species comparative framework. Population genetic analyses using 8-17 microsatellite markers per species (80 markers in total) as well as Bayesian inference were performed to estimate population structure and contemporary migration rates to highlight connectivity patterns and directionality of gene flow between depths. The results revealed three main genetic patterns depending on the species: a horizontal stepping stone pattern between areas around the island, a vertical stepping stone pattern between adjacent depths, and a quasi-panmictic pattern. Each species showed some specificity within these patterns, but overall, at least 4 of the 7 species support the assumption of vertical connectivity from the Deep Reef Refuge Hypothesis, highlighting the importance of studying multiple species. The existence of vertical connectivity between euphotic and mesophotic depths in the southwestern Indian Ocean confirms the importance of mesophotic coral ecosystems for conservation efforts and our global understanding of coral reef ecosystem dynamics.

Animals

Multistage Genetic, Transcriptomic, and Single-Cell Evidence Prioritizes MAP1LC3A among Ferroptosis-Related Genes in Glioblastoma.

Glioblastoma (GBM) remains a highly aggressive malignancy, and the contribution of ferroptosis-related genes to disease susceptibility remains incompletely understood. A genetically anchored, multistage framework was applied to prioritize ferroptosis-related genes associated with GBM. Among 483 genes curated from FerrDb V2, 315 had candidate cis-expression quantitative trait loci (cis-eQTLs) in eQTLGen, 250 retained at least three independent instruments after linkage disequilibrium clumping, and 226 yielded valid inverse-variance weighted (IVW) Mendelian randomization estimates using a GBM genome-wide association study comprising 6,183 cases and 18,169 controls. Thirty-four genes met the exploratory discovery criteria of P < 0.05 and a Benjamini-Hochberg false discovery rate (BH-FDR) < 0.20, with directionally concordant Bayesian weighted Mendelian randomization (BWMR) estimates. Replication-stage Mendelian randomization using GTEx V10 whole-blood cis-eQTLs supported four genes: ATG7, RPTOR, MAP1LC3A, and CHMP6. Evaluation across three independent tumor-control transcriptomic cohorts demonstrated that MAP1LC3A was consistently downregulated in tumor tissue and showed a significant random-effects pooled estimate (log&#x2082; fold change, -1.273; 95% confidence interval, -1.625 to -0.920; false discovery rate = 0.016), whereas the other three genes lacked comparable cross-cohort statistical support. Single-cell virtual knockout analysis was subsequently performed in a patient-balanced subset of 2,400 malignant cells selected from 4,916 eligible cells across 20 adult IDH-wild-type GBM tumors. Across five independently seeded runs, 3, 15, 4, and 7 robust downstream genes were identified for ATG7, RPTOR, MAP1LC3A, and CHMP6, respectively. The resulting consensus sets comprised 17 unique genes, with RND3 shared across all four targets. Gene Ontology analysis indicated enrichment of cell-adhesion and cell-surface processes, whereas no KEGG or Reactome pathways remained significant after multiple-testing correction. Collectively, these findings prioritize MAP1LC3A for future experimental investigation while distinguishing genetic association, tumor-expression concordance, and computational perturbation from definitive evidence of causality or mechanism.

Humans

Estimating Re and overdispersion in secondary cases from the size of identical sequence clusters of SARS-CoV-2.

The wealth of genomic data that was generated during the COVID-19 pandemic provides an exceptional opportunity to obtain information on the transmission of SARS-CoV-2. Specifically, there is great interest to better understand how the effective reproduction number [Formula: see text] and the overdispersion of secondary cases, which can be quantified by the negative binomial dispersion parameter k, changed over time and across regions and viral variants. The aim of our study was to develop a Bayesian framework to infer [Formula: see text] and k from viral sequence data. First, we developed a mathematical model for the distribution of the size of identical sequence clusters, in which we integrated viral transmission, the mutation rate of the virus, and incomplete case-detection. Second, we implemented this model within a Bayesian inference framework, allowing the estimation of [Formula: see text] and k from genomic data only. We validated this model in a simulation study. Third, we identified clusters of identical sequences in all SARS-CoV-2 sequences in 2021 from Switzerland, Denmark, and Germany that were available on GISAID. We obtained monthly estimates of the posterior distribution of [Formula: see text] and k, with the resulting [Formula: see text] estimates slightly lower than estimates obtained by other methods, and k comparable with previous results. We found comparatively higher estimates of k in Denmark which suggests less opportunities for superspreading and more controlled transmission compared to the other countries in 2021. Our model included an estimation of the case detection and sampling probability, but the estimates obtained had large uncertainty, reflecting the difficulty of estimating these parameters simultaneously. Our study presents a novel method to infer information on the transmission of infectious diseases and its heterogeneity using genomic data. With increasing availability of sequences of pathogens in the future, we expect that our method has the potential to provide new insights into the transmission and the overdispersion in secondary cases of other pathogens.

COVID-19

Estimating protein isoform abundances with [Formula: see text].

A single gene can encode multiple versions of a protein, dubbed isoforms, with varying functionality. Cellular control of isoform abundances is critical for multiple aspects of biology and is only partially regulated by transcript levels. While long-read sequencing facilitates transcript quantification, quantifying the resulting protein isoforms on a large scale is a major challenge, complicating biological interpretation of transcript alterations. Standard "bottom up" mass spectrometry can assess only short portions of isoforms called peptides, and these peptides often map onto more than one isoform. We introduce [Formula: see text] (Protein isoform Abundance Quantification), a Bayesian method that leverages multiomic information from the peptidome and transcriptome to provide accurate estimates of isoform abundance even when peptide mapping is ambiguous. [Formula: see text] offers several advantages over existing methods in a unified framework. It provides uncertainty quantification, integrates multiomic information for improved accuracy, and provides a rigorous framework for hypothesis testing. Extensive simulations show that [Formula: see text] consistently outperforms competing methods in detecting differentially abundant protein isoforms and estimating their abundances. We use [Formula: see text] to investigate differences in isoform abundance levels between people with schizophrenia and control subjects, confirming a long-held hypothesis that levels of the C4A isoform of Complement Component 4 are increased in schizophrenia while C4B is not. These results demonstrate that [Formula: see text] can identify significant variations in isoform abundance levels not previously possible.

Protein Isoforms

Genomic background of gestation length and calving-related traits in Holstein cattle.

The reproductive success of cows directly influences the profitability of dairy farms. Reproductive traits, particularly calving-related traits, generally have low heritability but sufficient additive genetic variance to enable genetic progress through genomic selection. Thus, the primary objectives of this study were to estimate genetic parameters and perform single-step genome-wide association studies (ssGWAS) for calf size, calving ease, gestation length, and stillbirth in Holstein cattle. Variance components were estimated based on animal models and Bayesian inference using a data set containing 226,717 animals with phenotypic records, 15,761 animals genotyped with 45,101 SNP markers, and 461,819 animals in the pedigree. SNP effects were estimated using the single-step GBLUP method. For direct and maternal genetic effects, heritability estimates (posterior standard deviation) ranged from 0.001 (0.002) for gestation length in heifers to 0.16 (0.001) for gestation length in cows. Genetic correlations ranged from -0.57 (0.01) between calving ease and stillbirth in heifers to 0.74 (0.01) between gestation length evaluated in heifers and cows. The ssGWAS results supported a highly polygenic architecture for calving-related traits, with most genomic signals not reaching genome-wide significance. A genome-wide significant association was detected for calving ease in cows on BTA23, highlighting FARS2 as a positional candidate gene. The strongest GWAS signals for each trait harbored additional biologically important candidate genes, including NPPA, NPPB, BCHE, EPHA4, DLD, and GTF2I. Given the generally low heritability estimates and the predominantly polygenic architecture observed for these traits, genomic selection may contribute to the genetic improvement of calving-related traits in Holstein cattle, with potential benefits for cow welfare, calf survival, and overall dairy production efficiency.

dairy cattle

Temporal reconstruction of a Salmonella Enteritidis ST11 outbreak in New Zealand.

Outbreaks caused by Salmonella Enteritidis are commonly linked to eggs and poultry meat internationally, but this serovar had never been detected in Aotearoa New Zealand (NZ) poultry prior to 2021. Locally designated genomic cluster Salmonella Enteritidis_2019_C_01, was implicated in a 2019 outbreak associated with a restaurant in Auckland. Four Enteritidis_2019_C_01 sub-clusters have since been identified, two retrospectively, in the Auckland region. Authorities initiated a formal outbreak investigation after genomically indistinguishable S. Enteritidis was isolated from the NZ poultry production environment. This study analysed 231&#x2009;S. Enteritidis genomes obtained from the outbreak using Bayesian phylodynamic tools to gain insight into the outbreak's dynamics and origin. We used Bayesian integrated coalescent epoch plots to estimate the change of the Enteritidis ST11 population size over time and marginal structured coalescent approximation to estimate transmission between poultry producers. We investigated human and poultry isolates to elucidate the time and location of the most recent common ancestor of the outbreak and transmission pathways. The median most recent common ancestor was estimated to be February 2019. We found evidence of amplification and spread of strain Enteritidis_2019_C_01 within the poultry industry, as well as transmission events throughout the production chain. The intervention by the public health and food safety authorities coincided with a drop in the effective population size of the S. Enteritidis ST11 as well as notified human cases. This information is crucial for understanding and preventing the transmission of S. Enteritidis in NZ poultry to ensure poultry meat and eggs are safe for consumption.

Salmonella enteritidis