Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Estimating Re and overdispersion in secondary cases from the size of identical sequence clusters of SARS-CoV-2.

The wealth of genomic data that was generated during the COVID-19 pandemic provides an exceptional opportunity to obtain information on the transmission of SARS-CoV-2. Specifically, there is great interest to better understand how the effective reproduction number [Formula: see text] and the overdispersion of secondary cases, which can be quantified by the negative binomial dispersion parameter k, changed over time and across regions and viral variants. The aim of our study was to develop a Bayesian framework to infer [Formula: see text] and k from viral sequence data. First, we developed a mathematical model for the distribution of the size of identical sequence clusters, in which we integrated viral transmission, the mutation rate of the virus, and incomplete case-detection. Second, we implemented this model within a Bayesian inference framework, allowing the estimation of [Formula: see text] and k from genomic data only. We validated this model in a simulation study. Third, we identified clusters of identical sequences in all SARS-CoV-2 sequences in 2021 from Switzerland, Denmark, and Germany that were available on GISAID. We obtained monthly estimates of the posterior distribution of [Formula: see text] and k, with the resulting [Formula: see text] estimates slightly lower than estimates obtained by other methods, and k comparable with previous results. We found comparatively higher estimates of k in Denmark which suggests less opportunities for superspreading and more controlled transmission compared to the other countries in 2021. Our model included an estimation of the case detection and sampling probability, but the estimates obtained had large uncertainty, reflecting the difficulty of estimating these parameters simultaneously. Our study presents a novel method to infer information on the transmission of infectious diseases and its heterogeneity using genomic data. With increasing availability of sequences of pathogens in the future, we expect that our method has the potential to provide new insights into the transmission and the overdispersion in secondary cases of other pathogens.

COVID-19↗

Comparison of phylogenetic metrics of transmission between symptomatic and asymptomatic tuberculosis in individuals who were incarcerated in Brazil in 2008-24: a retrospective genomic epidemiology study.

BACKGROUND: Tuberculosis control efforts have traditionally targeted symptomatic individuals; however, the role of asymptomatic cases in sustaining transmission is increasingly recognised. We aimed to quantify the contribution of asymptomatic tuberculosis to recent transmission using genomic and epidemiological data from a high-transmission setting. METHODS: We conducted a retrospective genomic epidemiology study of Mycobacterium tuberculosis isolates collected in Mato Grosso do Sul, Brazil, between Aug 25, 2008, and March 19, 2024. Available isolates underwent whole-genome sequencing. Demographic, clinical, incarceration history, and laboratory metadata were obtained from surveillance records. From Jan 1, 2017, to March 19, 2024, active case finding was conducted in the state's three largest prisons (all male-only facilities), during which sputum samples were collected from individuals irrespective of symptoms and tested using GeneXpert and culture. Comparisons of transmission between individuals with and without symptoms were restricted to individuals who were incarcerated and were identified through active case finding and for whom high-quality, M tuberculosis lineage 4 genomes were available. Metrics of recent transmission included phylogenetic clustering, time-scaled haplotype density (THD), local branching index (LBI), and transmission probabilities inferred using Bayesian Reconstruction and Evolutionary Analysis of Transmission Histories. FINDINGS: 4448 tuberculosis cases were notified in Mato Grosso do Sul in 2008-24. After excluding cases for which M tuberculosis isolates were not available or had low sequencing quality, who had contaminated cultures or mixed infection, or who were infected with non-lineage 4 M tuberculosis, we included 2362 lineage 4 M tuberculosis isolates with high-quality genome sequences. 1849 (78·3%) of 2362 isolates were part of a genomic cluster. Among 2362 individuals with tuberculosis, 1137 (48·1%) were incarcerated at diagnosis. Of these individuals, 505 were identified through active case finding in three male-only prisons. The median age was 30 years (IQR 25-37); 304 (60·2%) had mixed ethnicity, 90 (17·8%) were White, 56 (11·1%) were Black, 13 (2·6%) were Indigenous, and six (1·2%) were Asian. 277 (54·9%) had symptomatic disease and 228 (45·1%) had asymptomatic tuberculosis. There were no significant differences between symptomatic and asymptomatic individuals in phylogenetic clustering (213 [76·9%] of 277 vs 195 [85·5%] of 228; p=0·37), THD (median 0·39 [IQR 0·06-0·62] vs 0·50 [0·09-0·65]; p=0·12), or LBI (0·00863 [0·00810-0·00988] vs 0·00871 [0·00829-0·01020]; p=0·088). Bayesian transmission trees showed no significant difference in the number of secondary infections inferred from symptomatic compared with asymptomatic individuals (p=0·56). These findings were consistent across genomic clusters and robust to model assumptions. INTERPRETATION: We identified no differences in transmission between individuals who were symptomatic and those who were asymptomatic using multiple genomic measures. In this high-transmission setting, where systematic screening is implemented, our findings indicate that asymptomatic tuberculosis substantially contributes to tuberculosis transmission at the population level. These results suggest that symptom-based case detection alone is likely to be insufficient to interrupt transmission and highlight the importance of expanded screening strategies in high-risk populations. FUNDING: US National Institutes of Health and the Brazilian National Research Council (CNPq).

Humans↗

Should physicians be bayesian agents?

Because physicians use scientific inference for the generalizations of individual observations and the application of general knowledge to particular situations, the Bayesian probability solution to the problem of induction has been proposed and frequently utilized. Several problems with the Bayesian approach are introduced and discussed. These include: subjectivity, the favoring of a weak hypothesis, the problem of the false hypothesis, the old evidence/new theory problem and the observation that physicians are not currently Bayesians. To the complaint that the prior probability is subjective, Bayesians reply that there will be ultimate convergence, but the rebuttal to this is that there will not be uniform convergence. Secondly, since the Bayesian scheme favors a weak hypothesis, theories turn out to be a gratuitous risk. The problem with the false hypothesis comes out in the denominator of the theorem, revealing that a factor which is not a theory at all is being considered in the reasoning. On the old evidence/new theory problem old evidence cannot confirm a new theory so that the posterior probability will equal the prior probability. Finally, empiric studies have shown that current physicians are not Bayesians. But on consideration of Bayesian inference as a system of inference, it can be reasoned that physicians should be Bayesians. However, the problem of physicians' and patients' own subjectivity continue to plague this system of medical decision making.

Bayes Theorem↗

[Inferring genotype of DNA molecular marker by Bayesian theorem].

Bayesian theorem is applied to infer DNA molecular marker genotype (DNA chain type) from its phenotype (electrophoresis band type). The results indicate that large difference often presents in the genotype probability of a molecular marker with incomplete genetic information when it is obtained from the assumption of independence among markers as compared with that inferred from the genotypes of the flanking markers with the complete genetic information and the recombination fractions among them based on the Bayesian theorem. Therefore, before utilizing the marker information, such as in mapping quantitative trait loci (QTL), marker assisted selection (MAS) etc., Bayes' probability of the genotype for all markers with incomplete genetic information must be calculated over the whole genome for every individual. This study provided detailed procedure for the calculation of the Bayes' probability of the unknown DNA genotype. Several extensions were also discussed for the application of the Bayesian theorem in genetics.

Bayes Theorem↗

A Bayesian regression approach to the inference of regulatory networks from gene expression data.

MOTIVATION: There is currently much interest in reverse-engineering regulatory relationships between genes from microarray expression data. We propose a new algorithmic method for inferring such interactions between genes using data from gene knockout experiments. The algorithm we use is the Sparse Bayesian regression algorithm of Tipping and Faul. This method is highly suited to this problem as it does not require the data to be discretized, overcomes the need for an explicit topology search and, most importantly, requires no heuristic thresholding of the discovered connections. RESULTS: Using simulated expression data, we are able to show that this algorithm outperforms a recently published correlation-based approach. Crucially, it does this without the need to set any ad hoc threshold on possible connections.

Algorithms↗

A Bayesian analysis of regression models with continuous errors with application to longitudinal studies.

We employ a regression model with errors that follow a continuous autoregressive process to analyse longitudinal studies. In this way, unequally spaced observations do not present a problem in the analysis. We employ a Bayesian approach, where our inferences are based on a direct resampling process that generates values from the posterior distribution of the parameters of the model. We illustrate these Bayesian inferences with an analysis of a longitudinal study that involves the regression of foetal head circumference on menstrual age. Using these same data, we contrast the Bayesian approach with a maximum likelihood technique.

Bayes Theorem↗

Bayesian phylogenetic analysis of combined data.

The recent development of Bayesian phylogenetic inference using Markov chain Monte Carlo (MCMC) techniques has facilitated the exploration of parameter-rich evolutionary models. At the same time, stochastic models have become more realistic (and complex) and have been extended to new types of data, such as morphology. Based on this foundation, we developed a Bayesian MCMC approach to the analysis of combined data sets and explored its utility in inferring relationships among gall wasps based on data from morphology and four genes (nuclear and mitochondrial, ribosomal and protein coding). Examined models range in complexity from those recognizing only a morphological and a molecular partition to those having complex substitution models with independent parameters for each gene. Bayesian MCMC analysis deals efficiently with complex models: convergence occurs faster and more predictably for complex models, mixing is adequate for all parameters even under very complex models, and the parameter update cycle is virtually unaffected by model partitioning across sites. Morphology contributed only 5% of the characters in the data set but nevertheless influenced the combined-data tree, supporting the utility of morphological data in multigene analyses. We used Bayesian criteria (Bayes factors) to show that process heterogeneity across data partitions is a significant model component, although not as important as among-site rate variation. More complex evolutionary models are associated with more topological uncertainty and less conflict between morphology and molecules. Bayes factors sometimes favor simpler models over considerably more parameter-rich models, but the best model overall is also the most complex and Bayes factors do not support exclusion of apparently weak parameters from this model. Thus, Bayes factors appear to be useful for selecting among complex models, but it is still unclear whether their use strikes a reasonable balance between model complexity and error in parameter estimates.

Animals↗

Phylogenetic utility of protein (RPB2, beta-tubulin) and ribosomal (LSU, SSU) gene sequences in the systematics of Sordariomycetes (Ascomycota, Fungi).

The Sordariomycetes is an important group of fungi whose taxonomic relationships and classification is obscure. There is presently no multi-gene molecular phylogeny that addresses evolutionary relationships among different classes and orders. In this study, phylogenetic analyses with a broad taxon sampling of the Sordariomycetes were conducted to evaluate the utility of four gene regions (LSU rDNA, SSU rDNA, beta-tubulin and RPB2) for inferring evolutionary relationships at different taxonomic ranks. Single and multi-gene genealogies inferred from Bayesian and Maximum Parsimony analyses were compared in individual and combined datasets. At the subclass level, SSU rDNA phylogenies demonstrate their utility as a marker to infer phylogenetic relationships at higher levels. All analyses with SSU rDNA alone, combined LSU rDNA and SSU rDNA, and the combined 28 S rDNA, SSU rDNA and RPB2 datasets resulted in three subclasses: Hypocreomycetidae, Sordariomycetidae and Xylariomycetidae, which correspond well to established morphological classification schemes. At the ordinal level, the best resolved phylogeny was obtained from the combined LSU rDNA and SSU rDNA datasets. Individually, the RPB2 gene dataset resulted in significantly higher number of parsimony informative characters. Our results supported the recent separation of Boliniaceae, Chaetosphaeriaceae and Coniochaetaceae from Sordariales and placement of Coronophorales in Hypocreomycetidae. Microascales was found to be paraphyletic and Ceratocystis is phylogenetically associated to Faurelina, while Microascus and Petriella formed another clade and basal to other members of Halosphaeriales. In addition, the order Lulworthiales does not appear to fit in any of the three subclasses. Congruence between morphological and molecular classification schemes is discussed.

Ascomycota↗

Using complexity for the estimation of Bayesian networks.

Statistical inference of graphical models has become an important tool in the reconstruction of biological networks of the type which model, for example, gene regulatory interactions. In particular, the construction of a score-based Bayesian posterior density over the space of models provides an intuitive and computationally feasible method of assessing model uncertainty and of assigning statistical confidence to structural features. One problem which frequently occurs with this approach is the tendency to overestimate the degree of model complexity. Spurious graphical features obtained in this way may affect the inference in unpredictable ways, even when using scoring techniques, such as the Bayesian Information Criterion (BIC), that are specifically designed to compensate for overfitting. In this article we propose a simple adjustment to a BIC-based scoring procedure. The method proceeds in two steps. In the first step we derive an independent estimate of the parametric complexity of the model. In the second we modify the BIC score so that the mean parametric complexity of the posterior density is equal to the estimated value. The method is applied to a set of test networks, and to a collection of genes from the yeast genome known to possess regulatory relationships. A Bayesian network model with binary responses is employed. In the examples considered, we find that the number of spurious graph edges inferred is reduced, while the effect on the identification of true edges is minimal.

Algorithms↗

Influence of network topology and data collection on network inference.

We recently developed an approach for testing the accuracy of network inference algorithms by applying them to biologically realistic simulations with known network topology. Here, we seek to determine the degree to which the network topology and data sampling regime influence the ability of our Bayesian network inference algorithm, NETWORKINFERENCE, to recover gene regulatory networks. NETWORKINFERENCE performed well at recovering feedback loops and multiple targets of a regulator with small amounts of data, but required more data to recover multiple regulators of a gene. When collecting the same number of data samples at different intervals from the system, the best recovery was produced by sampling intervals long enough such that sampling covered propagation of regulation through the network but not so long such that intervals missed internal dynamics. These results further elucidate the possibilities and limitations of network inference based on biological data.

Algorithms↗

Evaluation of decay times in coupled spaces: Bayesian parameter estimation.

Determination of sound decay times in coupled spaces often demands considerable effort. Based on Schroeder's backward integration of room impulse responses, it is often difficult to distinguish different portions of multirate sound energy decay functions. A model-based parameter estimation method, using Bayesian probabilistic inference, proves to be a powerful tool for evaluating decay times. A decay model due to one of the authors [N. Xiang, J. Acoust. Soc. Am. 98, 2112-2121 (1995)] is extended to multirate decay functions. Following a summary of Bayesian model-based parameter estimation, the present paper discusses estimates in terms of both synthesized and measured decay functions. No careful estimation of initial values is required, in contrast to gradient-based approaches. The resulting robust algorithmic estimation of more than one decay time, from experimentally measured decay functions, is clearly superior to the existing nonlinear regression approach.

Journal Article↗

TreeFlow: Probabilistic Modelling and Automatic Differentiation for Phylogenetics.

Probabilistic modelling frameworks are powerful tools for statistical modelling and inference. They are not immediately generalizable to phylogenetic problems due to the particular computational properties of the phylogenetic tree object. TreeFlow is a software library for probabilistic modelling and automatic differentiation with phylogenetic trees. It embeds phylogenetic trees in the TensorFlow Probability framework, and implements inference algorithms for phylogenetic models given a fixed tree topology. We demonstrate how TreeFlow can be used to quickly implement and assess new models. We also show that it provides reasonable performance for gradient-based inference algorithms compared to specialized computational libraries for phylogenetics.

Bayesian inference↗

Bayesian approach to searching for susceptibility genes: Gc2 and EsD1 alleles and multiple sclerosis.

Multiple sclerosis (MS) is one of the most common causes of neurological disability in early adulthood. The current literature is interested in identifying biological or DNA markers associated with genetic susceptibility to MS. The aim of this study is to investigate, by means of Bayesian statistical inference, whether the presence of Gc2 (Gc = group-specific component) and/or EsD1 (EsD = esterase D) alleles affects MS susceptibility. Gc and EsD are two classical genetic markers, being the first a serum protein polymorphism, the latter an isoenzyme polymorphism. The interest of the proposed statistical approach of searching for MS susceptibility genes relies on the analysis of two different functions, one function being inferred from our results on 56 unrelated patients from central Italy affected by MS, the other one from Italian and worldwide epidemiological data. The graphical analysis suggests that MS susceptibility is influenced by both Gc2 and EsD1 alleles; and EsD1 allele is more informative than Gc2. These results point out the advantages of the Bayesian approach in searching for susceptibility genes. Furthermore, the significant association between the considered alleles and the susceptibility to MS suggests possible hypotheses about the pathogenesis of the disease.

Bayes Theorem↗

Perils of paralogy: using HSP70 genes for inferring organismal phylogenies.

Conserved genes have found their way into the mainstream of molecular systematics. Many of these genes are members of multigene families. A difficulty with using single genes of multigene families for phylogenetic inference is that genes from one species may be paralogous to those from another taxon. We focus attention on this problem using heat shock 70 (HSP70) genes. Using polymerase chain reaction techniques with genomic DNA, we isolated and sequenced 123 distinct sequences from 12 species of sharks. Phylogenetic analysis indicated that the sequences cluster with constituitively expressed cytoplasmic heat shock-like genes. Three highly divergent gene clades were sampled. A number of similar sequences were sampled from each species within each distinct gene clade. Comparison of published species trees with an HSP70 gene tree inferred using Bayesian phylogenetic analysis revealed several cases of gene duplication and differential sorting of gene lineages within this group of sharks. Gene tree parsimony based on the objective criteria of duplication and losses showed that previously published hypotheses of species relationships and two novel hypothesis based on Bayesian phylogenetics were concordant with the history of HSP70 gene duplication and loss. By contrast, two published hypotheses based on morphological data were not significantly different from the null hypothesis of a random association between species relatedness and the HSP70 gene tree. These results suggest that gene tree parsimony using data from multigene families can be used for inferring species relationships or testing published alternative hypotheses. More importantly, the results suggest that systematic studies relying on phylogenetic inferences from HSP70 genes may by plagued by unrecognized paralogy of sampled genes. Our results underscore the distinction between gene and species trees and highlight an underappreciated source of discordance between gene trees and organismal phylogeny, i.e., unrecognized paralogy of sampled genes.

Animals↗

Large subunit mitochondrial rRNA secondary structures and site-specific rate variation in two lizard lineages.

A phylogenetic-comparative approach was used to assess and refine existing secondary structure models for a frequently studied region of the mitochondrial encoded large subunit (16S) rRNA in two large lizard lineages within the Scincomorpha, namely the Scincidae and the Lacertidae. Potential pairings and mutual information were analyzed to identify site interactions present within each lineage and provide consensus secondary structures. Many of the interactions proposed by previous models were supported, but several refinements were possible. The consensus structures allowed a detailed analysis of rRNA sequence evolution. Phylogenetic trees were inferred from Bayesian analyses of all sites, and the topologies used for maximum likelihood estimation of sequence evolution parameters. Assigning gamma-distributed relative rate categories to all interacting sites that were homologous between lineages revealed substantial differences between helices. In both lineages, sites within helix G2 were mostly conserved, while those within helix E18 evolved rapidly. Clear evidence of substantial site-specific rate variation (covarion-like evolution) was also detected, although this was not strongly associated with specific helices. This study, in conjunction with comparable findings on different, higher-level taxa, supports the ubiquitous nature of site-specific rate variation in this gene and justifies the incorporation of covarion models in phylogenetic inference.

Animals↗

Population structure within and between subspecies of the Mediterranean triplefin fish Tripterygion delaisi revealed by highly polymorphic microsatellite loci.

Although F(ST) values are widely used to elucidate population relationships, in some cases, when employing highly polymorphic loci, they should be regarded with caution, particularly when subspecies are under consideration. Tripterygion delaisi presents two subspecies that were investigated here, using 10 microsatellite loci. A Bayesian approach allowed us to clearly identify both subspecies as two different evolutionary significant units. However, low F(ST) values were found between subspecies as a consequence of the large number of alleles per locus, while homoplasy could be disregarded as indicated by the standardized genetic distance G'(ST). Heterozygosity saturation was observed in highly polymorphic loci containing more than 15 alleles, and this threshold was used to define two loci pools. The less variable loci pool revealed higher genetic variance between subspecies, while the more variable pool showed higher genetic variance between populations. Furthermore, higher differentiation was also observed between populations using G'(ST) with the more variable loci. Nonetheless, a more reliable population structure within subspecies was obtained when all loci were included in the analyses. In T. d. xanthosoma, isolation by distance was detected between the eight analysed populations, and six genetically homogeneous clusters were inferred by Bayesian analyses that are in accordance with F(ST) values. The neighbourhood-size method also indicated rather small dispersal capabilities. In conclusion, in fish with limited adult and larval dispersal capabilities, continuous rocky habitat seems to allow contact between populations and prevent genetic differentiation, while large discontinuities of sand or deep-water channels seems to reduce gene flow.

Alleles↗

Inference network-based analyses of the histopathological effects of androgen deprivation on prostate cancer.

The evaluation of prostate cancer histology following hormonal therapy often represents a diagnostic problem for the pathologist. Previous studies have shown that an inference or Bayesian belief network (BBN) offers a descriptive classifier useful for the accurate analysis of morphological changes in individual cases of prostate neoplasia. Three different BBNs were evaluated in 94 cancer foci present in 20 radical prostatectomy (RP) specimens and in the matching biopsies in which the initial diagnosis of prostatic adenocarcinoma was made. Ten RP specimens were from patients treated with total androgen ablation or combination endocrine therapy (CET) before surgery. The first and second BBN allowed the identification with high certainty of the cancer foci present in the biopsies and RP specimens, as well as their Gleason grade, the belief value often being close to 1.0. The results of the second BBN showed a good correspondence between the Gleason grade given in the biopsies and that in the RP specimens, except in the surgical material of the treated patients, in which upgrading was always present. The third BBN showed the existence of three subgroups in treated RP specimens, one with morphological effect, another with poor effect, and the third with the histology of untreated (i.e. unaffected) cancer. In conclusion, an inference network-based analysis allows the characterization of treated prostate cancers according to the degree of histopathological change.

Adenocarcinoma↗

Hierarchical Bayesian spatial modelling of small-area rates of non-rare disease.

We present Bayesian hierarchical spatial models for the analysis of the geographical distribution of a non-rare disease or event. The work is motivated by the need for ascertaining regional variations in health services outcomes and resource use and for assessing the potential sources of these variations. The models discussed herein readily accommodate random spatial effects and covariate effects. We discuss Bayesian inferential framework and implementation of a hybrid Markov chain Monte Carlo method for full Bayesian model inference. The methods are illustrated through an analysis of regional variation in chronic lung disease (CLD) rates among neonatal intensive care unit (NICU) patients across Canada. Specifically, we first present a random effects binomial model for spatially correlated CLD rates, with random spatial effects accounting for latent or covariate effects. These random spatial effects depict regional or spatial variation in chronic lung disease occurrence. We then extend this model to include covariates. With this extension, we assess residual spatial effects and the extent to which risk factors such as illness severity at NICU admission, low birth weight, and very low birth weight influence the CLD rate variation.

Bayes Theorem↗