Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Inferring the sensitivity of wastewater metagenomic sequencing for early detection of viruses: a statistical modelling study.

BACKGROUND: Metagenomic sequencing of wastewater (W-MGS) can in principle detect any known or novel pathogen in a population. We aimed to quantify the sensitivity and cost of W-MGS for viral pathogen detection by jointly analysing W-MGS and epidemiological data for a range of human-infecting viruses. METHODS: In this statistical modelling study, we analysed sequencing data from four studies of untargeted W-MGS to estimate the relative abundance of 11 human-infecting viruses. Corresponding prevalence and incidence estimates were obtained or calculated from academic and public health reports. We combined these estimates using a hierarchical Bayesian model to predict relative abundance at set prevalence or incidence values, allowing comparison across studies and viruses. These predictions were then used to estimate the sequencing depth and concomitant cost required for pathogen detection using W-MGS with or without use of a hybridisation capture enrichment panel. FINDINGS: After controlling for variation in local infection rates, relative abundance varied by orders of magnitude across studies for a given virus. For instance, a local SARS-CoV-2 weekly incidence of 1% corresponded to a predicted SARS-CoV-2 relative abundance ranging from 3·8 × 10-10 to 2·4 × 10-7 across studies, translating to orders-of-magnitude variation in the cost of operating a system able to detect a SARS-CoV-2-like pathogen at a given sensitivity. Use of a respiratory virus enrichment panel in two studies greatly increased predicted relative abundance of SARS-CoV-2, lowering yearly costs by 27-fold (from US$7·87 million to $287 000) and 29-fold (from $1·98 million to $69 100) for a system able to detect a SARS-CoV-2-like pathogen before reaching 0·01% cumulative incidence. INTERPRETATION: The large variation in viral relative abundance after controlling for epidemiological factors indicates that other sources of inter-study variation, such as differences in sewershed hydrology and laboratory protocols, have a substantial impact on the sensitivity and cost of W-MGS. Well chosen hybridisation capture panels can greatly increase sensitivity and reduce cost for viruses in the panel, but might reduce sensitivity to unknown or unexpected pathogens. FUNDING: The Wellcome Trust, Open Philanthropy, and Musk Foundation.

Humans↗

Recent origin and phylogenetic utility of divergent ITS putative pseudogenes: a case study from Naucleeae (Rubiaceae).

The internal transcribed spacer (ITS) of nuclear ribosomal DNA has been widely used by systematists for reconstructing phylogenies of closely related taxa. Although the occurrence of ITS putative pseudogenes is well documented for many groups of animals and plants, the potential utility of these pseudogenes in phylogenetic analyses has often been underestimated or even ignored in part because of deletions that make unambiguous alignment difficult. In addition, long branches often can lead to spurious relationships, particularly in parsimony analyses. We have discovered unusually high levels of ITS polymorphism (up to 30%, 40%, and 14%, respectively) in three tropical tree species of the coffee family (Rubiaceae), Adinauclea fagifolia, Haldina cordifolia, and Mitragyna rubrostipulata. Both secondary structure stability and patterns of nucleotide substitutions in a highly conserved region (5.8S gene) were used for distinguishing presumed functional sequences from putative pseudogenes. The combination of both criteria was the most powerful approach. The sequences from A. fagifolia appear to be a mix of functional genes and highly distinct putative pseudogenes, whereas those from H. cordifolia and M. rubrostipulata were identified as putative pseudogenes. We explored the potential utility of the identified putative pseudogenes in the phylogenetic analyses of Naucleeae sensu lato. Both Bayesian and parsimony trees identified the same monophyletic groups and indicated that the polymorphisms do not transcend species boundaries, implying that they do not predate the divergence of these three species. The resulting trees are similar to those produced by previous analyses of chloroplast genes. In contrast to results of previous studies therefore, divergent putative pseudogenes can be useful for phylogenetic analyses, especially when no sequences of their functional counterparts are available. Our studies clearly show that ITS polymorphism may not necessarily mislead phylogenetic inference. Despite using many different PCR conditions (different primers, higher denaturing temperatures, and absence or presence of DMSO and BSA-TMACl), we recovered only a few functional ITS copies from A. fagifolia and none from H. cordifolia and M. rubrostipulata, which suggests that PCR selection is occurring and/or the presumed functional alleles are located at minor loci (with few ribosomal DNA copies).

Base Composition↗

Molecular and morphological phylogenetics of weevils (coleoptera, curculionoidea): do niche shifts accompany diversification?

The main goals of this study were to provide a robust phylogeny for the families of the superfamily Curculionoidea, to discover relationships and major natural groups within the family Curculionidae, and to clarify the evolution of larval habits and host-plant associations in weevils to analyze their role in weevil diversification. Phylogenetic relationships among the weevils (Curculionoidea) were inferred from analysis of nucleotide sequences of 18S ribosomal DNA (rDNA; approximately 2,000 bases) and 115 morphological characters of larval and adult stages. A worldwide sample of 100 species was compiled to maximize representation of weevil morphological and ecological diversity. All families and the main subfamilies of Curculionoidea were represented. The family Curculionidae sensu lato was represented by about 80 species in 30 "subfamilies" of traditional classifications. Phylogenetic reconstruction was accomplished by parsimony analysis of separate and combined molecular and morphological data matrices and Bayesian analysis of the molecular data; tree topology support was evaluated. Results of the combined analysis of 18S rDNA and morphological data indicate that monophyly of and relationships among each of the weevil families are well supported with the topology ((Nemonychidae, Anthribidae) (Belidae (Attelabidae (Caridae (Brentidae, Curculionidae))))). Within the clade Curculionidae sensu lato, the basal positions are occupied by mostly monocot-associated taxa with the primitive type of male genitalia followed by the Curculionidae sensu stricto, which is made up of groups with the derived type of male genitalia. High support values were found for the monophyly of some distinct curculionid groups such as Dryophthorinae (several tribes represented) and Platypodinae (Tesserocerini plus Platypodini), among others. However, the subfamilial relationships in Curculionidae are unresolved or weakly supported. The phylogeny estimate based on combined 18S rDNA and morphological data suggests that diversification in weevils was accompanied by niche shifts in host-plant associations and larval habits. Pronounced conservatism is evident in larval feeding habits, particularly in the host tissue consumed. Multiple shifts to use of angiosperms in Curculionoidea were identified, each time associated with increases in weevil diversity and subsequent shifts back to gymnosperms, particularly in the Curculionidae.

Animals↗

Phylogeny of Agrodiaetus Hübner 1822 (Lepidoptera: Lycaenidae) inferred from mtDNA sequences of COI and COII and nuclear sequences of EF1-alpha: karyotype diversification and species radiation.

Butterflies in the large Palearctic genus Agrodiaetus (Lepidoptera: Lycaenidae) are extremely uniform and exhibit few distinguishing morphological characters. However, these insects are distinctive in one respect: as a group they possess among the greatest interspecific karyotype diversity in the animal kingdom, with chromosome numbers (n) ranging from 10 to 125. The monophyly of Agrodiaetus and its systematic position relative to other groups within the section Polyommatus have been controversial. Characters from the mitochondrial genes for cytochrome oxidases I and II and from the nuclear gene for elongation factor 1 alpha were used to reconstruct the phylogeny of Agrodiaetus using maximum parsimony and Bayesian phylogenetic methods. Ninety-one individuals, encompassing most of the taxonomic diversity of Agrodiaetus, and representatives of 14 related genera were included in this analysis. Our data indicate that Agrodiaetus is monophyletic. Representatives of the genus Polyommatus (sensu stricto) are the closest relatives. The sequences of the Agrodiaetus taxa in this analysis are tentatively arranged into 12 clades, only 1 of which corresponds to a species group traditionally recognized in Agrodiaetus. Heterogeneous substitution rates across a recovered topology were homogenized with a nonparametric rate-smoothing algorithm before the application of a molecular clock. Two published estimates of substitution rates dated the origin of Agrodiaetus between 2.51 and 3.85 million years ago. During this time, there was heterogeneity in the rate and direction of karyotype evolution among lineages within the genus. Karyotype instability has evolved independently three times in the section Polyommatus, within the lineages Agrodiaetus, Lysandra, and Plebicula. Rapid karyotype diversification may have played a significant role in the radiation of the genus Agrodiaetus.

Animals↗

Speciational history of Australian grass finches (Poephila) inferred from thirty gene trees.

Multilocus genealogical approaches are still uncommon in phylogeography and historical demography, fields which have been dominated by microsatellite markers and mitochondrial DNA, particularly for vertebrates. Using 30 newly developed anonymous nuclear loci, we estimated population divergence times and ancestral population sizes of three closely related species of Australian grass finches (Poephila) distributed across two barriers in northern Australia. We verified that substitution rates were generally constant both among lineages and among loci, and that intralocus recombination was uncommon in our dataset, thereby satisfying two assumptions of our multilocus analysis. The reconstructed gene trees exhibited all three possible tree topologies and displayed considerable variation in coalescent times, yet this information provided the raw data for maximum likelihood and Bayesian estimation of population divergence times and ancestral population sizes. Estimates of these parameters were in close agreement with each other regardless of statistical approach and our Bayesian estimates were robust to prior assumptions. Our results suggest that black-throated finches (Poephila cincta) diverged from long-tailed finches (P. acuticauda and P. hecki) across the Carpentarian Barrier in northeastern Australia around 0.6 million years ago (mya), and that P. acuticauda diverged from P. hecki across the Kimberley Plateau-Arnhem Land Barrier in northwestern Australia approximately 0.3 mya. Bayesian 95% credibility intervals around these estimates strongly support Pleistocene timing for both speciation events, despite the fact that many gene divergences across the Carpentarian region clearly predated the Pleistocene. Estimates of ancestral effective population sizes for the basal ancestor and long-tailed finch ancestor were large (about 521,000 and about 384,000, respectively). Although the errors around the population size parameter estimates are considerable, they are the first for birds taking into account multiple sources of variance.

Animals↗

Predictions from uncertain categorizations.

Eleven experiments investigated how categorization influences feature prediction. Subjects were provided with sets of categorized exemplars, which they used to make predictions about properties of new exemplars. Because the categories were provided for subjects, this method allowed a test of categorization and prediction processes, bypassing initial concept formation and memory. The experiments tested a Bayesian rule of prediction according to which (1) predictions of an object's features are based on information from multiple categories, and (2) features are treated as independent of one another. With one exception, the studies found evidence against both of these claims. Subjects did not generally alter their predictions as a function of information outside the most likely "target" category. In addition, feature relations had reliable effects on these predictions. We discuss the implications of these results for understanding how categories are used in drawing inferences.

Concept Formation↗

Resolving deep phylogenetic relationships in salamanders: analyses of mitochondrial and nuclear genomic data.

Phylogenetic relationships among salamander families illustrate analytical challenges inherent to inferring phylogenies in which terminal branches are temporally very long relative to internal branches. We present new mitochondrial DNA sequences, approximately 2,100 base pairs from the genes encoding ND1, ND2, COI, and the intervening tRNA genes for 34 species representing all 10 salamander families, to examine these relationships. Parsimony analysis of these mtDNA sequences supports monophyly of all families except Proteidae, but yields a tree largely unresolved with respect to interfamilial relationships and the phylogenetic positions of the proteid genera Necturus and Proteus. In contrast, Bayesian and maximum-likelihood analyses of the mtDNA data produce a topology concordant with phylogenetic results from nuclear-encoded rRNA sequences, and they statistically reject monophyly of the internally fertilizing salamanders, suborder Salamandroidea. Phylogenetic simulations based on our mitochondrial DNA sequences reveal that Bayesian analyses outperform parsimony in reconstructing short branches located deep in the phylogenetic history of a taxon. However, phylogenetic conflicts between our results and a recent analysis of nuclear RAG-1 gene sequences suggest that statistical rejection of a monophyletic Salamandroidea by Bayesian analyses of our mitochondrial genomic data is probably erroneous. Bayesian and likelihood-based analyses may overestimate phylogenetic precision when estimating short branches located deep in a phylogeny from data showing substitutional saturation; an analysis of nucleotide substitutions indicates that these methods may be overly sensitive to a relatively small number of sites that show substitutions judged uncommon by the favored evolutionary model.

Animals↗

Computational strategy for discovering druggable gene networks from genome-wide RNA expression profiles.

We propose a computational strategy for discovering gene networks affected by a chemical compound. Two kinds of DNA microarray data are assumed to be used: One dataset is short time-course data that measure responses of genes following an experimental treatment. The other dataset is obtained by several hundred single gene knock-downs. These two datasets provide three kinds of information; (i) A gene network is estimated from time-course data by the dynamic Bayesian network model, (ii) Relationships between the knocked-down genes and their regulatees are estimated directly from knock-down microarrays and (iii) A gene network can be estimated by gene knock-down data alone using the Bayesian network model. We propose a method that combines these three kinds of information to provide an accurate gene network that most strongly relates to the mode-of-action of the chemical compound in cells. This information plays an essential role in pharmacogenomics. We illustrate this method with an actual example where human endothelial cell gene networks were generated from a novel time course of gene expression following treatment with the drug fenofibrate, and from 270 novel gene knock-downs. Finally, we succeeded in inferring the gene network related to PPAR-alpha, which is a known target of fenofibrate.

Bayes Theorem↗

Robust Bayesian prediction of subject disease status and population prevalence using several similar diagnostic tests.

Sometimes several diagnostic tests are performed on the same population of subjects with the aim of assessing disease status of individuals and the prevalence of the disease in the population, but no test is a reference test. Although the diagnostic tests may have the same biological underpinnings, test results may disagree for some specific animals. In that case, it may be difficult to determine disease status for individual subjects, and consequently population prevalence estimation becomes difficult. In this paper, we propose a robust method of estimating disease status and prevalence that uses heavy-tailed sampling distributions in a hierarchical model to protect against the influence of conflicting observations on inferences. If a subject has a test outcome that is discordant with the other test results then it is downweighted in diagnosing a subject's disease status, and for estimating disease prevalence. The amount of downweighting depends on the degree of conflict among the test results for the subject.

Animals↗

Model-free deconvolution of femtosecond kinetic data.

Though shorter laser pulses can also be produced, pulses of the 100 fs range are typically used in femtosecond kinetic measurements, which are comparable to characteristic times of the studied processes, making detection of the kinetic response functions inevitably distorted by convolution with the pulses applied. A description of this convolution in terms of experiments and measurable signals is given, followed by a detailed discussion of a large number of available methods to solve the convolution equation to get the undistorted kinetic signal, without any presupposed kinetic or photophysical model of the underlying processes. A thorough numerical test of several deconvolution methods is described, and two iterative time-domain methods (Bayesian and Jansson deconvolution) along with two inverse filtering frequency-domain methods (adaptive Wiener filtering and regularization) are suggested to use for the deconvolution of experimental femtosecond kinetic data sets. Adaptation of these methods to typical kinetic curve shapes is described in detail. We find that the model-free deconvolution gives satisfactory results compared to the classical "reconvolution" method where the knowledge of the kinetic and photophysical mechanism is necessary to perform the deconvolution. In addition, a model-free deconvolution followed by a statistical inference of the parameters of a model function gives less biased results for the relevant parameters of the model than simple reconvolution. We have also analyzed real-life experimental data and found that the model-free deconvolution methods can be successfully used to get undistorted kinetic curves in that case as well. A graphical computer program to perform deconvolution via inverse filtering and additional noise filters is also provided as Supporting Information. Though deconvolution methods described here were optimized for femtosecond kinetic measurements, they can be used for any kind of convolved data where measured experimental shapes are similar.

Journal Article↗

Inferring population history from microsatellite and enzyme data in serially introduced cane toads, Bufo marinus.

Much progress has been made on inferring population history from molecular data. However, complex demographic scenarios have been considered rarely or have proved intractable. The serial introduction of the South-Central American cane toad Bufo marinus in various Caribbean and Pacific islands involves four major phases: a possible genetic admixture during the first introduction, a bottleneck associated with founding, a transitory population boom, and finally, a demographic stabilization. A large amount of historical and demographic information is available for those introductions and can be combined profitably with molecular data. We used a Bayesian approach to combine this information with microsatellite (10 loci) and enzyme (22 loci) data and used a rejection algorithm to simultaneously estimate the demographic parameters describing the four major phases of the introduction history. The general historical trends supported by microsatellites and enzymes were similar. However, there was a stronger support for a larger bottleneck at introductions for microsatellites than enzymes and for a more balanced genetic admixture for enzymes than for microsatellites. Very little information was obtained from either marker about the transitory population boom observed after each introduction. Possible explanations for differences in resolution of demographic events and discrepancies between results obtained with microsatellites and enzymes were explored. Limits of our model and method for the analysis of nonequilibrium populations were discussed.

Animals↗

Handling manipulated evidence.

Bayesian Networks have been advocated as useful tools to describe the relations of dependence/independence among random variables and relevant hypotheses in a crime case. Moreover, they have been applied to help the investigator structure the problem and evaluate the impact of the observed evidence, typically with respect to the hypothesis of guilt of a suspect. In this paper we describe a model to handle the possibility that one or more pieces of evidence have been manipulated in order to mislead the investigations. This method is based on causal inference models, although it is developed in a different, specific framework.

Bayes Theorem↗

The footprints of visual attention in the Posner cueing paradigm revealed by classification images.

In the Posner cueing paradigm, observers' performance in detecting a target is typically better in trials in which the target is present at the cued location than in trials in which the target appears at the uncued location. This effect can be explained in terms of a Bayesian observer where visual attention simply weights the information differently at the cued (attended) and uncued (unattended) locations without a change in the quality of processing at each location. Alternatively, it could also be explained in terms of visual attention changing the shape of the perceptual filter at the cued location. In this study, we use the classification image technique to compare the human perceptual filters at the cued and uncued locations in a contrast discrimination task. We did not find statistically significant differences between the shapes of the inferred perceptual filters across the two locations, nor did the observed differences account for the measured cueing effects in human observers. Instead, we found a difference in the magnitude of the classification images, supporting the idea that visual attention changes the weighting of information at the cued and uncued location, but does not change the quality of processing at each individual location.

Attention↗

Mixed Bayesian networks: a mixture of Gaussian distributions.

Mixed Bayesian networks are probabilistic models associated with a graphical representation, where the graph is directed and the random variables are discrete or continuous. We propose a comprehensive method for estimating the density functions of continuous variables, using a graph structure and a set of samples. The principle of the method is to learn the shape of densities from a sample of continuous variables. The densities are approximated by a mixture of Gaussian distributions. The estimation algorithm is a stochastic version of the Expectation Maximization algorithm (Stochastic EM algorithm). The inference algorithm corresponding to our model is a variant of junction three method, adapted to our specific case. The approach is illustrated by a simulated example from the domain of pharmacokinetics. Tests show that the true distributions seem sufficiently fitted for practical application.

Algorithms↗

Bayesian analysis of mutational spectra.

Studies that examine both the frequency of gene mutation and the pattern or spectrum of mutational changes can be used to identify chemical mutagens and to explore the molecular mechanisms of mutagenesis. In this article, we propose a Bayesian hierarchical modeling approach for the analysis of mutational spectra. We assume that the total number of independent mutations and the numbers of mutations falling into different response categories, defined by location within a gene and/or type of alteration, follow binomial and multinomial sampling distributions, respectively. We use prior distributions to summarize past information about the overall mutation frequency and the probabilities corresponding to the different mutational categories. These priors can be chosen on the basis of data from previous studies using an approach that accounts for heterogeneity among studies. Inferences about the overall mutation frequency, the proportions of mutations in each response category, and the category-specific mutation frequencies can be based on posterior distributions, which incorporate past and current data on the mutant frequency and on DNA sequence alterations. Methods are described for comparing groups and for assessing dose-related trends. We illustrate our approach using data from the literature.

Animals↗

The complete mitochondrial genome of the nudibranch Roboastra europaea (Mollusca: Gastropoda) supports the monophyly of opisthobranchs.

The complete nucleotide sequence (14,472 bp) of the mitochondrial genome of the nudibranch Roboastra europaea (Gastropoda: Opisthobranchia) was determined. This highly compact mitochondrial genome is nearly identical in gene organization to that found in opisthobranchs and pulmonates (Euthyneura) but not to that in prosobranchs (a paraphyletic group including the most basal lineages of gastropods). The newly determined mitochondrial genome differs only in the relative position of the trnC gene when compared with the mitochondrial genome of Pupa strigosa, the only opisthobranch mitochondrial genome sequenced so far. Pupa and Roboastra represent the most basal and derived lineages of opisthobranchs, respectively, and their mitochondrial genomes are more similar in sequence when compared with those of pulmonates. All phylogenetic analyses (maximum parsimony, minimum evolution, maximum likelihood, and Bayesian) based on the deduced amino acid sequences of all mitochondrial protein-coding genes supported the monophyly of opisthobranchs. These results are in agreement with the classical view that recognizes Opisthobranchia as a natural group and contradict recent phylogenetic studies of the group based on shorter sequence data sets. The monophyly of opisthobranchs was further confirmed when a fragment of 2,500 nucleotides including the mitochondrial cox1, rrnL, nad6, and nad5 genes was analyzed in several species representing five different orders of opisthobranchs with all common methods of phylogenetic inference. Within opisthobranchs, the polyphyly of cephalaspideans and the monophyly of nudibranchs were recovered. The evolution of mitochondrial tRNA rearrangements was analyzed using the cox1+rrnL+nad6+nad5 gene phylogeny. The relative position of the trnP gene between the trnA and nad6 genes was found to be a synapomorphy of opisthobranchs that supports their monophyly.

Amino Acid Sequence↗

Phylogeny of venus clams (Bivalvia: Venerinae) as inferred from nuclear and mitochondrial gene sequences.

Venerinae (Heterodonta: Veneridae) is a diverse, commercially important, and cosmopolitan marine bivalve subfamily. Recent workers synonymized it with the subfamily Chioninae, due to their overall morphological similarity. The use of traditional shell-based characters alone, however, is questionable for resolving phylogenetic relationships of this group. A phylogenetic study was carried out, based on nucleotide sequences of the mitochondrial large ribosomal subunit (16S), cytochrome oxidase subunit I (COI), and the nuclear protein-coding gene histone 3, to investigate the relationships and circumscription of Venerinae and the phylogenetic pattern of characters in this group. This study consists of a total of 55 taxa: 13 venerine genera, 24 chionine taxa, and 18 taxa of other venerid subfamilies. We analyzed the alignments using a Bayesian approach using Markov Chain Monte Carlo tree sampling and maximum parsimony methods. The resulting phylogenetic hypothesis suggests that Chioninae and Venerinae are actually discrete taxa, but that the circumscription suffered from misplacement of some genera. Our analysis showed that the former chionine genera Chamelea and Clausinella should be placed in Venerinae, as sister taxa to Venus. We re-analyzed morphological and anatomical features in light of the molecular data to describe monophyletic entities. Features of the hinge and internal shell as well as the degree of siphonal fusion are identified as characters to morphologically distinguish the two subfamilies. Of the three genes used in this study, only COI (commonly used as "barcoding" gene) posed substantial problems in obtaining sequence data from older museum material.

Animals↗

NExON-Bayes: a Bayesian approach to network estimation informed by ordinal covariates.

MOTIVATION: In heterogeneous disease settings, accounting for intrinsic sample variability is crucial for obtaining reliable and interpretable omic network estimates. However, most graphical model analyses of biomedical data assume homogeneous conditional dependence structures, potentially leading to misleading conclusions. To address this, we propose a joint Gaussian graphical model that leverages sample-level ordinal covariates (e.g. disease stage) to account for heterogeneity and improve the estimation of partial correlation structures. RESULTS: Our modelling framework, called NExON-Bayes, extends the graphical spike-and-slab framework to account for ordinal covariates, jointly estimating their relevance to the graph structure and leveraging them to improve the accuracy of network estimation. To scale to high-dimensional omic settings, we develop an efficient variational inference algorithm tailored to our model. Through simulations, we demonstrate that our method outperforms the vanilla graphical spike-and-slab (with no covariate information), as well as other state-of-the-art network approaches which exploit covariate information. Applying our method to reverse phase protein array data from patients diagnosed with stage I, II or III breast carcinoma, we estimate the behaviour of proteomic networks as cancer progresses. Our model provides insights not only through inspection of the estimated proteomic networks, but also of the estimated ordinal covariate dependencies of key groups of proteins within those networks, offering a comprehensive understanding of how biological pathways shift across disease stages. AVAILABILITY AND IMPLEMENTATION: A user-friendly R package for NExON-Bayes with tutorials is available on Github at github.com/jf687/NExON, and archived at https://doi.org/10.5281/zenodo.20312938. The source of the dataset used is cited in the relevant section.

Bayes Theorem↗