Search PubMedSearch

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Internal Bayesian precision modulates the neural representation of social attention: Disentangling implicit and explicit components via model-informed multivariate EEG analysis.

Social attention integrates sensory cues with high-level cognitive expectations, yet the generative mechanisms through which implicit orienting and explicit belief-driven modulation interact remain poorly understood. This ambiguity complicates the distinction between specialized social modules and domain-general attentional processes. We combined a dynamic cueing task with hierarchical Bayesian modeling and model-informed multivariate EEG decoding to address this. Behavioral results revealed a computational double dissociation: symbolic arrow cues elicited heterogeneous strategies, whereas averted gaze recruited a consistent, surprise-driven computational phenotype. At the neural level, time-resolved decoding and temporal generalization revealed a critical representational shift starting approximately 400 ms post-cue. Initial activity related to physical cue features was rapidly replaced by stable neural templates of predicted spatial intent. Crucially, topographical activation patterns showed that this intentional template, characterized by a lateralized temporo-occipital distribution, emerged exclusively under high internal certainty. Furthermore, partial representational similarity analysis demonstrated that late-stage neural manifolds were overwhelmingly organized around integrated spatial goals rather than isolated sensory or motivational signals. These findings suggest that social attention is a specialized generative process, where internal certainty modulates the transformation of social perceptions into actionable top-down intentions.

Bayesian computational modeling

HiCPotts: An R/Bioconductor package to identify significant interactions in chromosome conformation capture data and model sources of bias.

MOTIVATION: Chromosome Conformation Capture methods, including Hi-C, micro-C or Capture-C, are used to map chromatin interactions genome-wide. Most of the existing computational methods do not account for sources of bias (such as DNA accessibility, GC content or TE content) in the data. RESULTS: We previously developed ZipHiC, a Bayesian method based on the hidden Markov random field (HMRF) model and the Approximate Bayesian Computation (ABC), that uses zero-inflated Poisson distribution to model the noise, signal and false signal of the data and showed that this approach was able to detect bias from DNA accessibility, GC content and TE content in both Hi-C and micro-C data. Here, we present HiCPotts, another Bayesian method based on the HMRF model and the ABC that uses a zero-inflated Negative Binomial distribution instead to model the noise and signal of the data. We systematically show that HiCPotts reduces false positives and increases recovery of true interactions compared to ZipHiC, but also compared to other methods such as FastHiC, Juicer and HiCExplorer. Most importantly, we provide an R/Bioconductor package that allows modelling the noise, signal and false signal using various distributions such as the zero-inflated Negative Binomial (ZINB) and the zero-inflated Poisson distribution (ZIP). AVAILABILITY AND IMPLEMENTATION: https://bioconductor.org/packages/HiCPotts/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Approximate Bayesian Computation

Modelling the effects of biological intervention in a dynamical gene network.

Cellular response to environmental and internal signals can be modeled by dynamical gene regulatory networks (GRN). In the literature, three main classes of gene network models can be distinguished: (1) non-quantitative (or data-based) models which do not describe the probability distribution of gene expressions; (2) quantitative models which fully describe the probability distribution of all genes co-expression; and (3) mechanistic models which allow for a causal interpretation of gene interactions. We propose two rigorous frameworks to model gene alteration in a dynamical GRN, depending on whether the network model is quantitative or mechanistic. We explain how these models can be used for design of experiment, or, if additional alteration data are available, for validation purposes or to improve the parameter estimation of the original model. We apply these methods to the Gaussian graphical model, which is quantitative but non-mechanistic, and to mechanistic models of Bayesian networks and penalized linear regression.

Gene Regulatory Networks

ScITree: Scalable Bayesian inference of transmission tree from epidemiological and genomic data.

Phylodynamic models capture joint epidemiological-evolutionary dynamics during an outbreak, providing a powerful tool to enhance understanding and management of disease transmission. Existing phylodynamic approaches, however, mostly rely on various non-mechanistic or semi-mechanistic approximations of the underlying epidemiological-evolutionary process. Previous work by Lau and colleagues has shown that full Bayesian mechanistic models, without relying on these approximations, can enable highly accurate joint inference of the epidemiological-evolutionary dynamics including the unobserved transmission tree. However, the Lau method faces major computational bottlenecks. As the volume of genomic data collected during outbreaks continues to grow, it is crucial to develop scalable yet accurate phylodynamic methods. Here we propose a new Bayesian phylodynamic model, overcoming the major scalability issue in the previous method and enabling a readily deployable, yet accurate, phylodynamic modeling framework. Specifically, we develop a scalable spatio-temporal phylodynamic framework for inferring the transmission tree (ScITree) and other key epidemiological parameters considering the infinite sites assumption in modeling mutation on the sequence level, in contrast to the Lau method in which mutation was modeled explicitly on the nucleotide level. Our approach features full Bayesian implementation utilizing an exact likelihood to mechanistically integrate epidemiological and evolutionary processes. We develop a computationally-efficient data-augmentation Markov Chain Monte Carlo algorithm, inferring key model parameters and unobserved dynamics including the transmission tree. We assess performance of our method using multiple simulated outbreak datasets. Our results indicate that our method can achieve high inference accuracy, comparable to the performance of the Lau method. Additionally, our method scales significantly more efficiently for large outbreaks, with computing time increasing linearly with outbreak size, compared to the exponential scaling of the Lau method. We also demonstrate our method's utility by applying our validated modeling framework to a dataset describing a foot-and-mouth disease outbreak in the UK. Our results show that our method is able to generate estimates of the transmission dynamics consistent with those from the prior method, further demonstrating the robustness of our new approach. In summary, our method provides a computationally-efficient, highly scalable, accurate modeling framework for inferring the joint spatio-temporal dynamics of epidemiological and evolutionary processes, facilitating timely and effective outbreak responses in space and time. Our method is implemented in our R package ScITree.

Bayes Theorem

Bayesian reconstruction and differential testing of excised introns.

MOTIVATION: Characterizing the differential excision of introns is critical for understanding the functional complexity of a cell or tissue, from normal developmental processes to disease pathogenesis. Most transcript reconstruction methods infer full-length transcripts from high-throughput sequencing data. However, this is a challenging task due to incomplete annotations and the heterogeneous expression of transcripts across cell-types, tissues, and experimental conditions. Several recent methods circumvent these difficulties by considering local splicing events, but these methods lose transcript-level splicing information and may conflate similar, but distinct transcripts. RESULTS: In this work, we formalize a new transcript reconstruction problem that interpolates between the full-length and local splicing perspectives by considering sequences of exon-exon junctions (SEEJs) that co-occur in transcripts. We then present a hierarchical Bayesian admixture model and posterior inference algorithms for computing SEEJs (BSEEJ), and a generalized linear model for characterizing differential SEEJ usage based on model parameter estimates. We show that BSEEJ achieves high F1 score for reconstruction tasks and improved accuracy and sensitivity in differential splicing when compared with six transcript and local splicing methods on simulated data. Lastly, we evaluate BSEEJ on experimental data based on transcript reconstruction, novelty of transcripts produced, model sensitivity to hyperparameters, and a functional analysis of differentially expressed SEEJs. AVAILABILITY AND IMPLEMENTATION: BSEEJ is freely available at https://github.com/bayesomicslab/BSEEJ.

Bayes Theorem

Mutational signatures in blood-brain barrier: mechanisms, computational insights, and clinical applications in precision oncology.

The blood - brain barrier (BBB) plays a central role in maintaining central nervous system (CNS) homeostasis, and its disruption is a defining feature of malignant brain tumors such as glioblastoma. Emerging evidence indicates that BBB dysfunction not only alters the tumor microenvironment but also shapes the mutational processes that drive genomic instability in CNS malignancies. This review synthesizes current understanding of the biological mechanisms linking BBB breakdown with distinct mutational signatures, including those arising from oxidative stress, hypoxia-induced replication stress, lipid peroxidation, inflammation, and metabolic reprogramming. Advances in next-generation sequencing, coupled with computational tools such as non-negative matrix factorization, Bayesian modeling, and deep learning, have enabled precise extraction of these signatures and their integration with multi-omics data. Clinically, BBB-associated mutational signatures offer significant promise for therapeutic stratification, prediction of treatment response, and noninvasive monitoring through cerebrospinal fluid - derived circulating tumor DNA. Despite these advances, challenges persist due to limited tissue accessibility, low-yield CSF samples, incomplete mechanistic models, and the lack of CNS-specific analytical frameworks. A deeper understanding of BBB-driven mutational processes, supported by improved computational approaches and integrative datasets, holds potential to advance precision oncology in neuro-oncology.

Humans

Bayesian Modeling of Cancer Outcomes Using Genetic Variables Assisted by Pathological Imaging Data.

With the increasing maturity of genetic profiling, an essential and routine task in cancer research is to model disease outcomes/phenotypes using genetic variables. Many methods have been successfully developed. However, oftentimes, empirical performance is unsatisfactory because of a "lack of information." In cancer research and clinical practice, a source of information that is broadly available and highly cost-effective comes from pathological images, which are routinely collected for definitive diagnosis and staging. In this article, we consider a Bayesian approach for selecting relevant genetic variables and modeling their relationships with a cancer outcome/phenotype. We propose borrowing information from (manually curated, low-dimensional) pathological imaging features via reinforcing the same selection results for the cancer outcome and imaging features. We further develop a weighting strategy to accommodate the scenario where information borrowing may not be equally effective for all subjects. Computation is carefully examined. Simulations demonstrate competitive performance of the proposed approach. We analyze TCGA (The Cancer Genome Atlas) LUAD (lung adenocarcinoma) data, with overall survival and gene expressions being the outcome and genetic variables, respectively. Findings different from the alternatives and with sound properties are made.

Humans

Non-destructive prediction of lead content in oilseed rape leaves by fluorescence hyperspectral technology based on neural network.

Based on fluorescence hyperspectral imaging (FHSI), this study targeted rapid, non-destructive quantification of lead (Pb) content in oilseed rape leaves treated with varying silicon (Si) concentrations, acquiring fluorescence spectra over the 484.43-1001.61 nm wavelength range. To optimize spectral data quality, preprocessing methods (Savitzky-Golay smoothing, first derivative, detrending) were comprehensively compared. Characteristic wavelengths were then selected via interval variable iterative shrinkage, which effectively compressed data dimensionality and reduced computational load. A hybrid SE-CL1DA model, fusing a 1D convolutional neural network, a long short-term memory network and SE attention mechanism was constructed, with Bayesian optimization tuning hyperparameters to boost stability. The BO-SE-CL1DA outperformed both traditional machine learning and insufficiently optimized deep learning model (Rp2=0.9609, RMSE = 0.0377 mg/kg, RPD = 5.1736), thus enabling accurate Pb estimation, supporting Si-regulated heavy metal stress management and facilitating agricultural contamination monitoring.

Plant Leaves

Towards improved fine-mapping of candidate causal variants.

Fine-mapping in genome-wide association studies aims to identify potentially causal genetic variants among a set of candidate variants that are often highly correlated with each other owing to linkage disequilibrium. A variety of statistical approaches are used in fine-mapping, almost all of which are based on a multiple regression framework to model the relationship between genotype and phenotype, while accommodating specific assumptions about the distribution of variant effect sizes and using different inference algorithms. Owing to their modelling flexibility and the ease of making inferential statements, these approaches are predominantly Bayesian in nature. Recently, these approaches have been improved by refining modelling assumptions, integrating additional information, accommodating summary statistics, and developing scalable computational algorithms that improve computation efficiency and fine-mapping resolution.

Humans

Parallel algorithms for phylogenetic inference under a structured coalescent approximation.

While advances in molecular epidemiology and computational modeling have enhanced our capacity to track pathogen evolution, the accurate reconstruction of spatiotemporal transmission dynamics remains essential for developing epidemic preparedness frameworks and implementing outbreak response measures. Structured coalescent models offer a phylogeographic framework by restricting lineage coalescence events to geographically proximate host populations. Although the Bayesian structured coalescent approximation (BASTA) provides a tractable approach, contemporary phylogeographic analyses involving dozens of geographic localities and hundreds to thousands of viral genomes substantially exceed the computational capacity of existing implementations. The BASTA likelihood scales cubically with deme count and quadratically with sequence count due to matrix exponentiation and pairwise coalescent probability calculations. Here, we introduce a comprehensive algorithmic restructuring of the structured coalescent likelihood that eliminates redundancies, optimizes memory access, and exposes parallelization opportunities. Our approach reorganizes computations along three dimensions: (i) independent calculation of deme-transition probability matrices across time intervals; (ii) simultaneous evaluation of partial likelihood vectors within temporal slices; and (iii) concurrent aggregation of coalescent probabilities. Algorithmic restructuring cuts average coalescent likelihood computation by 7-8 fold, and parallelization further boosts performance to 10-26 fold, enabling joint phylogeographic analyses of dengue virus across 10 South American countries and H5N1 avian influenza across 20 Eurasian regions to finish in a fraction of prior time. This computational efficiency also enables comparison between backward-in-time structured coalescent approximations and forward-in-time phylogeographic methods, revealing that the former provides appropriately conservative posterior estimates, particularly at intermediate phylogenetic depths. We integrate our implementation into the popular BEAST X and BEAGLE software packages, with an accompanying interface in BEAUti X to easily set up the analyses, providing researchers with an accessible and scalable tool for real-time phylogeographic surveillance of rapidly evolving pathogens.

Journal Article

Bayesian Inference of Pathogen Phylogeography using the Structured Coalescent Model.

Over the past decade, pathogen genome sequencing has become well established as a powerful approach to study infectious disease epidemiology. In particular, when multiple genomes are available from several geographical locations, comparing them is informative about the relative size of the local pathogen populations as well as past migration rates and events between locations. The structured coalescent model has a long history of being used as the underlying process for such phylogeographic analysis. However, the computational cost of using this model does not scale well to the large number of genomes frequently analysed in pathogen genomic epidemiology studies. Several approximations of the structured coalescent model have been proposed, but their effects are difficult to predict. Here we show how the exact structured coalescent model can be used to analyse a precomputed dated phylogeny, in order to perform Bayesian inference on the past migration history, the effective population sizes in each location, and the directed migration rates from any location to another. We describe an efficient reversible jump Markov Chain Monte Carlo scheme which is implemented in a new R package StructCoalescent. We use simulations to demonstrate the scalability and correctness of our method and to compare it with existing software. We also applied our new method to several state-of-the-art datasets on the population structure of real pathogens to showcase the relevance of our method to current data scales and research questions.

Bayes Theorem

TreeFlow: Probabilistic Modelling and Automatic Differentiation for Phylogenetics.

Probabilistic modelling frameworks are powerful tools for statistical modelling and inference. They are not immediately generalizable to phylogenetic problems due to the particular computational properties of the phylogenetic tree object. TreeFlow is a software library for probabilistic modelling and automatic differentiation with phylogenetic trees. It embeds phylogenetic trees in the TensorFlow Probability framework, and implements inference algorithms for phylogenetic models given a fixed tree topology. We demonstrate how TreeFlow can be used to quickly implement and assess new models. We also show that it provides reasonable performance for gradient-based inference algorithms compared to specialized computational libraries for phylogenetics.

Bayesian inference

Comparing ARG Inference Methods Under Transmission of Reproductive Success: Tree Imbalance Matters.

Inferring coalescent trees from genomic data has become a major subject in population genetics, particularly with the recent advances in tree sequence reconstruction methods. However, it remains unclear how well these methods perform for imbalanced genealogies. Such imbalances can arise from processes such as cultural transmission of reproductive success (CTRS) or positive selection. Using simulated genomic data, we benchmarked three major software packages, SINGER, Relate, and tsinfer, by comparing the imbalance of reconstructed trees by these methods with that of the true simulated trees, for three indices that quantify this imbalance. The three methods performed well under scenarios yielding balanced trees. However, their accuracy declined as imbalance increased. Performances also varied with mutation rate, recombination rate, and sample size. This study opens possibilities for applying these methods to infer CTRS or positive selection in large-scale genomic datasets, using simulation-based inference such as approximate Bayesian computation.

Models, Genetic

Diagnosing scientific replicability through probabilistic distinguishability.

MOTIVATION: Despite the widely recognized importance of replicability in biological research, computational methods to quantify irreplicability and identify irreplicable instances remain underdeveloped. This article presents an efficient and robust computational framework to address this gap. RESULTS: To tackle the challenge of defining an acceptable level of intrinsic heterogeneity among replicable studies, we introduce a distinguishability criterion, ensuring that replicable effects, while potentially heterogeneous, can be distinguished from zero effects and maintain consistent directions with high probability. We implement a Bayesian model criticism approach, reporting a Bayesian P-value to identify potential irreplicable instances. Through numerical experiments, we demonstrate the efficacy of the proposed methods in detecting batch effects in high-throughput experiments and identifying instances of the publication bias. Finally, we apply the framework to multi-tissue eQTL data from the GTEx consortium, uncovering tissue-specific eQTLs that represent biological heterogeneity across tissues. AVAILABILITY AND IMPLEMENTATION: An R package DiscRep implementing our method is available on GitHub (https://github.com/PengWang96/DiscRep).

Bayes Theorem

The signed two-space proximity model for learning representations in protein-protein interaction networks.

MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.

Protein Interaction Mapping

A Bayesian framework for multivariate differential analysis.

Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.

Bayes Theorem

Bayesian classification of OXPHOS deficient skeletal myofibres.

Mitochondria are organelles in most human cells which release the energy required for cells to function. Oxidative phosphorylation (OXPHOS) is a key biochemical process within mitochondria required for energy production and requires a range of proteins and protein complexes. Mitochondria contain multiple copies of their own genome (mtDNA), which codes for some of the proteins and ribonucleic acids required for mitochondrial function and assembly. Pathology arises from genetic defects in mtDNA and can reduce cellular abundance of OXPHOS proteins, affecting mitochondrial function. Due to the continuous turn-over of mtDNA, pathology is random and neighbouring cells can possess different OXPHOS protein abundance. Estimating the proportion of cells where OXPHOS protein abundance is too low to maintain normal function is critical to understanding disease severity and predicting disease progression. Currently, one method to classify single cells as being OXPHOS deficient is prevalent in the literature. The method compares a patient's OXPHOS protein abundance to that of a small number of healthy control subjects. If the patient's cell displays an abundance which differs from the abundance of the controls then it is deemed deficient. However, due to the natural variation between subjects and the low number of control subjects typically available, this method is inflexible and often results in a large proportion of patient cells being misclassified. These misclassifications have significant consequences for the clinical interpretation of these data. We propose a single-cell classification method using a Bayesian hierarchical mixture model, which allows for inter-subject OXPHOS protein abundance variation. The model accurately classifies an example dataset of OXPHOS protein abundances in skeletal muscle fibres (myofibres). When comparing the proposed and existing model classifications to manual classifications performed by experts, the proposed model results in estimates of the proportion of deficient myofibres that are consistent with expert manual classifications.

Oxidative Phosphorylation

TPMM: three-component posterior mixture model enables robust inverton detection in low-depth metagenomes and suggests potential viral invertons.

SUMMARY: Bacterial phase variation enables reversible, locus-specific phenotypic switching, often driven by DNA inversion (invertons). To identify these events, researchers commonly rely on sequencing reads that provide orientation-specific support. Metagenomic sequencing, which captures total genetic material independent of cultivation, offers a powerful platform for the comprehensive study of invertons. However, computational inverton calling from metagenomic data is difficult at low sequencing depth: hard read-support cutoffs can miss true events, while sequence-only predictors lack read-backed interpretability and uncertainty quantification. To address this, we present TPMM, a three-component posterior mixture model for inverton calling in metagenomic data. TPMM explicitly incorporates sequencing depth to formulate inverton detection as a probabilistic mixture problem. Starting from candidates flanked by inverted repeats, the model classifies the candidates into noise, low-probability, or high-probability inversion signals using read evidence. Finally, TPMM assigns posterior probabilities as soft labels and applies cumulative Bayesian False Discovery Rate control to robustly identify true invertons. On two real gut metagenomic datasets, TPMM agrees well with PhaseFinder at high depth but recovers substantially more invertons under systematic downsampling, demonstrating superior performance in sparse-data regimes. We further examine potential reversible inversion elements in viral genomes and provide supporting analyses, suggesting a broader scope for inversion-mediated regulation. AVAILABILITY: The source code of TPMM is available via: https://github.com/KennyxxD/TPMM.

Metagenomics