Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

Bayesian cost-effectiveness analysis. An example using the GUSTO trial.

A desirable element of cost-effectiveness analysis (CEA) modeling is a systematic way to relate uncertainty about input parameters to uncertainty in the computational results of the CEA model. Use of Bayesian statistical estimation and Monte Carlo simulation provides a natural way to compute a posterior probability distribution for each CEA result. We demonstrate this approach by reanalyzing a previously published CEA evaluating the incremental cost-effectiveness of tissue plasminogen activator compared to streptokinase for thrombolysis in acute myocardial infarction patients using data from the GUSTO trial and other auxiliary data sources. We illustrate Bayesian estimation for proportions, mean costs, and mean quality-of-life weights. The computations are performed using the Bayesian analysis software WinBUGS, distributed by the MRC Biostatistics Unit, Cambridge, England.

Bayes Theorem↗

Estimating the rate of evolution of the rate of molecular evolution.

A simple model for the evolution of the rate of molecular evolution is presented. With a Bayesian approach, this model can serve as the basis for estimating dates of important evolutionary events even in the absence of the assumption of constant rates among evolutionary lineages. The method can be used in conjunction with any of the widely used models for nucleotide substitution or amino acid replacement. It is illustrated by analyzing a data set of rbcL protein sequences.

Algorithms↗

Dispersal, vicariance, and timing of diversification in Nothonotus darters.

The species diversity of North American freshwater fishes is unparalleled among temperate regions of the planet. This diversity is concentrated in the Central Highlands of eastern North America and this distribution pattern has inspired different models involving either dispersal or vicariance to explain the high species diversity of North American fishes. The most popular of these models is the Central Highlands vicariance hypothesis (CHVH), which proposes an ancient and diverse widespread fauna that existed across a previously continuous highland landscape that is much different from today. The mechanisms of isolation in the CHVH involve specific instances of vicariance that affected several diverse lineages of Central Highlands fishes. We tested predictions of the CHVH and alternative models using a cytochrome b-inferred phylogeny of the darter clade Nothonotus. A Bayesian mixed-model method was used for phylogenetic analysis. The phylogenetic data set included all 20 recognized Nothonotus species, and most species were represented with multiple sequences. We were able to convert genetic branch lengths to absolute age using external fossil calibrations in the freshwater perciform fish clade Centrarchidae. Using a well-resolved Nothonotus phylogeny and divergence time estimates, we identify equal numbers of instances of both vicariance and dispersal among disjunct regions of the Central Highlands, biogeographic pseudocongruence, rather recent speciation in Nothonotus, and a surprisingly large amount of speciation within highland areas. With regard to Nothonotus, previous Central Highlands biogeographic models offer little in the way of providing possible mechanisms responsible for diversification in the clade. Patterns of speciation in Nothonotus are similar to those discovered in recent efforts that have included speciation as a parameter into classic models of island biogeography.

Animals↗

Bayesian analysis of stochastic constraints in structural equation models.

Structural equation models are analysed in the presence of stochastic constraints. Based on a Bayesian perspective, a prior distribution on nuisance parameters in the unknown covariance matrix of error measurements with stochastic constraints is considered. An iterative procedure is implemented to produce the various Bayesian estimates with stochastic constraints. A simulation study is conducted to illustrate the accuracy and behaviour of this Bayesian approach. A real-life example is provided to illustrate the theory.

Bayes Theorem↗

Modeling a Poisson forest in variable elevations: a nonparametric Bayesian approach.

A nonparametric Bayesian formulation is given to the problem of modeling nonhomogeneous spatial point patterns influenced by concomitant variables. Only incomplete information on the concomitant variables is assumed, consisting of a relatively small number of point measurements. Residual variation, caused by other unmeasured influential factors, is modeled in terms of a spatially varying baseline intensity function. A Markov chain Monte Carlo scheme is proposed for the simultaneous nonparametric estimation of each unknown function in the model. The suggested method is illustrated by reanalysing a data set in Rathbun (1996, Biometrics 52, 226-242), and the estimated models are compared with those obtained by Rathbun.

Altitude↗

Identifying protein complexes in high-throughput protein interaction screens using an infinite latent feature model.

We propose a Bayesian approach to identify protein complexes and their constituents from high-throughput protein-protein interaction screens. An infinite latent feature model that allows for multi-complex membership by individual proteins is coupled with a graph diffusion kernel that evaluates the likelihood of two proteins belonging to the same complex. Gibbs sampling is then used to infer a catalog of protein complexes from the interaction screen data. An advantage of this model is that it places no prior constraints on the number of complexes and automatically infers the number of significant complexes from the data. Validation results using affinity purification/mass spectrometry experimental data from yeast RNA-processing complexes indicate that our method is capable of partitioning the data in a biologically meaningful way. A supplementary web site containing larger versions of the figures is available at http://public.kgi.edu/wild/PSBO6/index.html.

Algorithms↗

Estimating Re and overdispersion in secondary cases from the size of identical sequence clusters of SARS-CoV-2.

The wealth of genomic data that was generated during the COVID-19 pandemic provides an exceptional opportunity to obtain information on the transmission of SARS-CoV-2. Specifically, there is great interest to better understand how the effective reproduction number [Formula: see text] and the overdispersion of secondary cases, which can be quantified by the negative binomial dispersion parameter k, changed over time and across regions and viral variants. The aim of our study was to develop a Bayesian framework to infer [Formula: see text] and k from viral sequence data. First, we developed a mathematical model for the distribution of the size of identical sequence clusters, in which we integrated viral transmission, the mutation rate of the virus, and incomplete case-detection. Second, we implemented this model within a Bayesian inference framework, allowing the estimation of [Formula: see text] and k from genomic data only. We validated this model in a simulation study. Third, we identified clusters of identical sequences in all SARS-CoV-2 sequences in 2021 from Switzerland, Denmark, and Germany that were available on GISAID. We obtained monthly estimates of the posterior distribution of [Formula: see text] and k, with the resulting [Formula: see text] estimates slightly lower than estimates obtained by other methods, and k comparable with previous results. We found comparatively higher estimates of k in Denmark which suggests less opportunities for superspreading and more controlled transmission compared to the other countries in 2021. Our model included an estimation of the case detection and sampling probability, but the estimates obtained had large uncertainty, reflecting the difficulty of estimating these parameters simultaneously. Our study presents a novel method to infer information on the transmission of infectious diseases and its heterogeneity using genomic data. With increasing availability of sequences of pathogens in the future, we expect that our method has the potential to provide new insights into the transmission and the overdispersion in secondary cases of other pathogens.

COVID-19↗

A Bayesian approach to generating tutorial hints in a collaborative medical problem-based learning system.

OBJECTIVES: Today a great many medical schools have turned to a problem-based learning (PBL) approach to teaching. While PBL has many strengths, effective PBL requires the tutor to provide a high degree of personal attention to the students, which is difficult in the current academic environment of increasing demands on faculty time. This paper describes intelligent tutoring in a collaborative medical tutor for PBL. The main contribution of our work is the development of representational techniques and algorithms for generating tutoring hints in PBL group problem solving, as well as the implementation of these techniques in a collaborative intelligent tutoring system, COMET. The system combines concepts from computer-supported collaborative learning with those from intelligent tutoring systems. METHODS AND MATERIALS: The system uses Bayesian networks to model individual student clinical reasoning, as well as that of the group. The prototype system incorporates substantial domain knowledge in the areas of head injury, stroke and heart attack. Tutoring in PBL is particularly challenging since the tutor should provide as little guidance as possible while at the same time not allowing the students to get lost. From studies of PBL sessions at a local medical school, we have identified and implemented eight commonly used hinting strategies. In order to evaluate the appropriateness and quality of the hints generated by our system, we compared the tutoring hints generated by COMET with those of experienced human tutors. We also compared the focus of group activity chosen by COMET with that chosen by human tutors. RESULTS: On average, 74.17% of the human tutors used the same hint as COMET. The most similar human tutor agreed with COMET 83% of the time and the least similar tutor agreed 62% of the time. Our results show that COMET's hints agree with the hints of the majority of the human tutors with a high degree of statistical agreement (McNemar test, p=0.652, kappa=0.773). The focus of group activity chosen by COMET agrees with that chosen by the majority of the human tutors with a high degree of statistical agreement (McNemar test, p=0.774, kappa=0.823). CONCLUSION: Bayesian network clinical reasoning models can be combined with generic tutoring strategies to successfully emulate human tutor hints in group medical PBL.

Artificial Intelligence↗

Multiple association analysis via simulated annealing (MASSA).

SUMMARY: Genome-wide association studies are now technically feasible and likely to become a fundamental tool in unraveling the ultimate genetic basis of complex traits. However, new statistical and computational methods need to be developed to extract the maximum information in a realistic computing time. Here we propose a new method for multiple association analysis via simulated annealing that allows for epistasis and any number of markers. It consists of finding the model with lowest Bayesian information criterion using simulated annealing. The data are described by means of a mixed model and new alternative models are proposed using a set of rules, e.g. new sites can be added (or deleted), or new epistatic interactions can be included between existing genetic factors. The method is illustrated with simulated and real data. AVAILABILITY: An executable version of the program (MASSA) running under the Linux OS is freely available, together with documentation, at http://www.icrea.es/pag.asp?id=Miguel.Perez.

Algorithms↗

Role of knowledge in human visual temporal integration in spatiotemporal noise.

Previous studies have shown how human observers' knowledge about the signal's spatial frequency, spatial phase, and spatial locations affects human performance in detecting and identifying signals in spatial noise. These results have led to the idea that human observers can be modeled as suboptimal Bayesian observers that use a priori information to generate probabilities or likelihoods for hypothesis. This approach has also been applied more recently to object recognition. We investigate whether human observers have the ability to use information about the temporal profile of a temporally modulated signal in temporal information processing. We measure human performance in detecting a time-varying signal embedded in spatiotemporal (dynamic) noise with and without a cue that contains information about the temporal phase of the signal. Results show improvement in performance in the phase-cued condition, suggesting that human observers act as if they have the ability to use knowledge about the temporal shape of the signal when performing temporal information processing. Human performance is consistent with a suboptimal Bayesian observer and a newly proposed Max-Min observer. The results also suggest that models based solely on the integration of the early temporal filters in the human visual system and/or any further integration (e.g., probability summation), which do not make use of knowledge about the signals' temporal profile, are incomplete models of human visual detection in spatiotemporal noise.

Adult↗

Inferring evolutionary signals from ecological data in a plant-pathogen metapopulation.

We followed the dynamics of local epidemics in three populations of a natural plant-pathogen system for four sequential years. We characterize the overwintering process with spatial statistics and use a stochastic, spatially explicit, modeling approach with Bayesian parameter estimation to study the spread of the infection during the growing season. Our modeling approach allows us to infer coevolutionary signals from spatiotemporal data on pathogen prevalence. Most importantly, we are able to assess the distribution of resistant hosts within the distribution of all host plants. We show that resistant hosts occur in areas with high pathogen encounter rates, and that the occurrence of resistance correlates with overwintering probability of the pathogen. The estimates for essentially all model parameters are characterized by a large amount of variation over the years and the populations. While the variation in the fraction of resistant hosts and in the force of infection is to a large extent explained by the population, the other model parameters (two parameters describing the shape of the dispersal kernel) vary essentially in an unpredictable manner, suggesting that much of the variation may occur at very fine spatial and temporal scales.

Bayes Theorem↗

Frequentist properties of Bayesian posterior probabilities of phylogenetic trees under simple and complex substitution models.

What does the posterior probability of a phylogenetic tree mean?This simulation study shows that Bayesian posterior probabilities have the meaning that is typically ascribed to them; the posterior probability of a tree is the probability that the tree is correct, assuming that the model is correct. At the same time, the Bayesian method can be sensitive to model misspecification, and the sensitivity of the Bayesian method appears to be greater than the sensitivity of the nonparametric bootstrap method (using maximum likelihood to estimate trees). Although the estimates of phylogeny obtained by use of the method of maximum likelihood or the Bayesian method are likely to be similar, the assessment of the uncertainty of inferred trees via either bootstrapping (for maximum likelihood estimates) or posterior probabilities (for Bayesian estimates) is not likely to be the same. We suggest that the Bayesian method be implemented with the most complex models of those currently available, as this should reduce the chance that the method will concentrate too much probability on too few trees.

Animals↗

Bayesian design criteria: computation, comparison, and application to a pharmacokinetic and a pharmacodynamic model.

In this paper 3 criteria to design experiments for Bayesian estimation of the parameters of nonlinear models with respect to their parameters, when a prior distribution is available, are presented: the determinant of the Bayesian information matrix, the determinant of the pre-posterior covariance matrix, and the expected information provided by an experiment. A procedure to simplify the computation of these criteria is proposed in the case of continuous prior distributions and is compared with the criterion obtained from a linearization of the model about the mean of the prior distribution for the parameters. This procedure is applied to two models commonly encountered in the area of pharmacokinetics and pharmacodynamics: the one-compartment open model with bolus intravenous single-dose injection and the Emax model. They both involve two parameters. Additive as well as multiplicative gaussian measurement errors are considered with normal prior distributions. Various combinations of the variances of the prior distribution and of the measurement error are studied. Our attention is restricted to designs with limited numbers of measurements (1 or 2 measurements). This situation often occurs in practice when Bayesian estimation is performed. The optimal Bayesian designs that result vary with the variances of the parameter distribution and with the measurement error. The two-point optimal designs sometimes differ from the D-optimal designs for the mean of the prior distribution and may consist of replicating measurements. For the studied cases, the determinant of the Bayesian information matrix and its linearized form lead to the same optimal designs. In some cases, the pre-posterior covariance matrix can be far from its lower bound, namely, the inverse of the Bayesian information matrix, especially for the Emax model and a multiplicative measurement error. The expected information provided by the experiment and the determinant of the pre-posterior covariance matrix generally lead to the same designs except for the Emax model and the multiplicative measurement error. Results show that these criteria can be easily computed and that they could be incorporated in modules for designing experiments.

Bayes Theorem↗

A segmentation-based regularization term for image deconvolution.

This paper proposes a new and original inhomogeneous restoration (deconvolution) model under the Bayesian framework for observed images degraded by space-invariant blur and additive Gaussian noise. In this model, regularization is achieved during the iterative restoration process with a segmentation-based a priori term. This adaptive edge-preserving regularization term applies a local smoothness constraint to pre-estimated constant-valued regions of the target image. These constant-valued regions (the segmentation map) of the target image are obtained from a preliminary Wiener deconvolution estimate. In order to estimate reliable segmentation maps, we have also adopted a Bayesian Markovian framework in which the regularized segmentations are estimated in the maximum a posteriori (MAP) sense with the joint use of local Potts prior and appropriate Gaussian conditional luminance distributions. In order to make these segmentations unsupervised, these likelihood distributions are estimated in the maximum likelihood sense. To compute the MAP estimate associated to the restoration, we use a simple steepest descent procedure resulting in an efficient iterative process converging to a globally optimal restoration. The experiments reported in this paper demonstrate that the discussed method performs competitively and sometimes better than the best existing state-of-the-art methods in benchmark tests.

Algorithms↗

A stochastic regression approach to analyzing thermodynamic uncertainty in chemical speciation modeling.

Chemical speciation modeling is a vital tool for assessing the bioavailability of inorganic species, yet significant uncertainties in thermodynamic parameters and model form limit its potential for decision-making. In this paper we present a novel method for the quantification of thermodynamic parameter uncertainty and ionic strength correction model uncertainty using Bayesian Markov Chain Monte Carlo (MCMC) estimation methods. These methods allow for the inclusion of correlation modeling, which has not been present in previous work. The MCMC simulations are used to model a natural river water to determine the uncertainty in the calculated environmental speciation of ethylenediamenetetraacetate, a chelating agent that has attracted considerable environmental interest. The results indicate that incorporating correlation among related thermodynamic parameters into the uncertainty model is necessary to correctly quantify the overall system uncertainty. This result indicates the superiority of MCMC estimation methods overtraditional Monte Carlo methods when available data are used to estimate parameter uncertainty in systems with closely related model parameters.

Bayes Theorem↗

Bayesian semiparametric analysis of developmental toxicology data.

Modeling of developmental toxicity studies often requires simple parametric analyses of the dose-response relationship between exposure and probability of a birth defect but poses challenges because of nonstandard distributions of birth defects for a fixed level of exposure. This article is motivated by two such experiments in which the distribution of the outcome variable is challenging to both the standard logistic model with binomial response and its parametric multistage elaborations. We approach our analysis using a Bayesian semiparametric model that we tailored specifically to developmental toxicology studies. It combines parametric dose-response relationships with a flexible nonparametric specification of the distribution of the response, obtained via a product of Dirichlet process mixtures approach (PDPM). Our formulation achieves three goals: (1) the distribution of the response is modeled in a general way, (2) the degree to which the distribution of the response adapts nonparametrically to the observations is driven by the data, and (3) the marginal posterior distribution of the parameters of interest is available in closed form. The logistic regression model, as well as many of its extensions such as the beta-binomial model and finite mixture models, are special cases. In the context of the two motivating examples and a simulated example, we provide model comparisons, illustrate overdispersion diagnostics that can assist model specification, show how to derive posterior distributions of the effective dose parameters and predictive distributions of response, and discuss the sensitivity of the results to the choice of the prior distribution.

2,4,5-Trichlorophenoxyacetic Acid↗

Bayesian network and nonparametric heteroscedastic regression for nonlinear modeling of genetic network.

We propose a new statistical method for constructing genetic network from microarray gene expression data by using a Bayesian network. An essential point of Bayesian network construction is in the estimation of the conditional distribution of each random variable. We consider fitting nonparametric regression models with heterogeneous error variances to the microarray gene expression data to capture the nonlinear structures between genes. A problem still remains to be solved in selecting an optimal graph, which gives the best representation of the system among genes. We theoretically derive a new graph selection criterion from Bayes approach in general situations. The proposed method includes previous methods based on Bayesian networks. We demonstrate the effectiveness of the proposed method through the analysis of Saccharomyces cerevisiae gene expression data newly obtained by disrupting 100 genes.

Artificial Intelligence↗

Bayesian analysis of twinning and ovulation rates using a multiple-trait threshold model and Gibbs sampling.

The Multiple-Trait Gibbs Sampler for Animal Models programs were extended to allow analysis of ordered categorical data using a Bayesian threshold model. The algorithm is based on data augmentation, where a value on the unobserved underlying normally distributed variable (liability) is generated in each round of iteration for each categorical observation. The programs allow analysis of several continuous and ordered categorical traits. Categorical traits can have any number of response levels. Models can be different for each trait. The programs were used to analyze twinning and ovulation rates from a herd of cattle selected for twinning rate at the U.S. Meat Animal Research Center. Data included number of calves born at each parturition for the lifetime of a cow and number of eggs ovulated for several estrous cycles before first breeding as heifers. A total of 6,411 calvings was recorded for 2,087 cows with 83.2% single and 16.8% multiple births. A total of 19,849 ovulations was recorded for 2,332 heifers with 85.2% single and 14.8% multiple ovulations. Mean posterior estimates of heritability and fraction of variance accounted for by permanent environmental effects (PE) were .128 and .103 for twinning rate and .168 and .079 for ovulation rate. Mean posterior estimate of genetic correlation was .808, and correlation of PE effects was .517. Use of a threshold model could allow for more rapid genetic improvement of the twinning herd through improved identification and selection of genetically superior animals because of higher heritability on the underlying scale.

Algorithms↗