Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Large-scale neural models and dynamic causal modelling.

Dynamic causal modelling (DCM) is a method for estimating and making inferences about the coupling among small numbers of brain areas, and the influence of experimental manipulations on that coupling [Friston, K.J., Harrison, L., Penny, W., 2003 Dynamic causal modelling. Neuroimage 19, 1273-1302]. Large-scale neural modelling aims to construct neurobiologically grounded computational models with emergent behaviours that inform our understanding of neuronal systems. One such model has been used to simulate region-specific BOLD time-series [Horwitz, B., Friston, K.J., Taylor, J.G., 2000. Neural modeling and functional brain imaging: an overview. Neural Netw. 13, 829-846]. DCM was used to make inferences about effective connectivity using data generated by a model implementing a visual delayed match-to-sample task [Tagamets, M.A., Horwitz, B., 1998. Integrating electrophysiological and anatomical experimental data to create a large-scale model that simulates a delayed match-to-sample human brain imaging study. Cereb. Cortex 8, 310-320]. The aim was to explore the validity of inferences made using DCM about the connectivity structure and task-dependent modulatory effects, in a system with a known connectivity structure. We also examined the effects of misspecifying regions of interest. Models with hierarchical connectivity and reciprocal connections were examined using DCM and Bayesian Model Comparison [Penny, W.D., Stephan, K.E., Mechelli, A., Friston, K.J., 2004. Comparing dynamic causal models. Neuroimage 22, 1157-1172]. This approach revealed strong evidence for those models with correctly specified anatomical connectivity. Furthermore, Bayesian model comparison favoured those models when bilinear effects corresponded to their implementation in the neural model. These findings generalised to an extended model with two additional areas and reentrant circuits. The conditional uncertainty of coupling parameter estimates increased in proportion to the number of incorrectly specified regions. These results highlight the role of neural models in establishing the validity of estimation and inference schemes. Specifically, Bayesian model comparison confirms the validity of DCM in relation to a well-characterised and comprehensive neuronal model.

Bayes Theorem↗

Feature inference and the causal structure of categories.

The purpose of this article was to establish how theoretical category knowledge-specifically, knowledge of the causal relations that link the features of categories-supports the ability to infer the presence of unobserved features. Our experiments were designed to test proposals that causal knowledge is represented psychologically as Bayesian networks. In five experiments we found that Bayes' nets generally predicted participants' feature inferences quite well. However, we also observed a pervasive violation of one of the defining principles of Bayes' nets-the causal Markov condition-because the presence of characteristic features invariably led participants to infer yet another characteristic feature. We argue that this effect arises from a domain-general bias to assume the presence of underlying mechanisms associated with the category. Specifically, people take an exemplar to be a "well functioning" category member when it has most or all of the category's characteristic features, and thus are likely to infer a characteristic value on an unobserved dimension.

Analysis of Variance↗

Lifting a veil on diversity: a Bayesian approach to fitting relative-abundance models.

Bayesian methods incorporate prior knowledge into a statistical analysis. This prior knowledge is usually restricted to assumptions regarding the form of probability distributions of the parameters of interest, leaving their values to be determined mainly through the data. Here we show how a Bayesian approach can be applied to the problem of drawing inference regarding species abundance distributions and comparing diversity indices between sites. The classic log series and the lognormal models of relative- abundance distribution are apparently quite different in form. The first is a sampling distribution while the other is a model of abundance of the underlying population. Bayesian methods help unite these two models in a common framework. Markov chain Monte Carlo simulation can be used to fit both distributions as small hierarchical models with shared common assumptions. Sampling error can be assumed to follow a Poisson distribution. Species not found in a sample, but suspected to be present in the region or community of interest, can be given zero abundance. This not only simplifies the process of model fitting, but also provides a convenient way of calculating confidence intervals for diversity indices. The method is especially useful when a comparison of species diversity between sites with different sample sizes is the key motivation behind the research. We illustrate the potential of the approach using data on fruit-feeding butterflies in southern Mexico. We conclude that, once all assumptions have been made transparent, a single data set may provide support for the belief that diversity is negatively affected by anthropogenic forest disturbance. Bayesian methods help to apply theory regarding the distribution of abundance in ecological communities to applied conservation.

Animals↗

Overcredibility of molecular phylogenies obtained by Bayesian phylogenetics.

Bayesian phylogenetics has recently been proposed as a powerful method for inferring molecular phylogenies, and it has been reported that the mammalian and some plant phylogenies were resolved by using this method. The statistical confidence of interior branches as judged by posterior probabilities in Bayesian analysis is generally higher than that as judged by bootstrap probabilities in maximum likelihood analysis, and this difference has been interpreted as an indication that bootstrap support may be too conservative. However, it is possible that the posterior probabilities are too high or too liberal instead. Here, we show by computer simulation that posterior probabilities in Bayesian analysis can be excessively liberal when concatenated gene sequences are used, whereas bootstrap probabilities in neighbor-joining and maximum likelihood analyses are generally slightly conservative. These results indicate that bootstrap probabilities are more suitable for assessing the reliability of phylogenetic trees than posterior probabilities and that the mammalian and plant phylogenies may not have been fully resolved.

Amino Acid Substitution↗

Diagnosis using predictive probabilities without cut-offs.

Standard diagnostic test procedures involve dichotomization of serologic test results. The critical value or cut-off is determined to optimize a trade off between sensitivity and specificity of the resulting test. When sampled units from a population are tested, they are allocated as either infected or not according to the test outcome. Units with values high above the cut-off are treated the same as units with values just barely above the cut-off, and similarly for values below the cut-off. There is an inherent information loss in dichotomization. We thus develop a diagnostic screening method based on data that are not dichotomized within the Bayesian paradigm. Our method determines the predictive probability of infection for each individual in a sample based on having observed a specific serologic test result and provides inferences about the prevalence of infection in the population sampled. Our fully Bayesian method is briefly compared with a previously developed frequentist method. We illustrate the methodology with serologic data that have been previously analysed in the veterinary literature, and also discuss applications to screening for disease in humans. The method applies more generally to a variation of the classic parametric 2-population discriminant analysis problem. Here, in addition to training data, additional units are sampled and the goal is to determine their population status, and the prevalence(s) of the subpopulation(s) from which they were sampled.

Animals↗

Hierarchical Bayesian modeling of spatially correlated health service outcome and utilization rates.

We present Bayesian hierarchical spatial models for spatially correlated small-area health service outcome and utilization rates, with a particular emphasis on the estimation of both measured and unmeasured or unknown covariate effects. This Bayesian hierarchical model framework enables simultaneous modeling of fixed covariate effects and random residual effects. The random effects are modeled via Bayesian prior specifications reflecting spatial heterogeneity globally and relative homogeneity among neighboring areas. The model inference is implemented using Markov chain Monte Carlo methods. Specifically, a hybrid Markov chain Monte Carlo algorithm (Neal, 1995, Bayesian Learning for Neural Networks; Gustafson, MacNab, and Wen, 2003, Statistics and Computing, to appear) is used for posterior sampling of the random effects. To illustrate relevant problems, methods, and techniques, we present an analysis of regional variation in intraventricular hemorrhage incidence rates among neonatal intensive care unit patients across Canada.

Bayes Theorem↗

Multiple imputation for model checking: completed-data plots with missing and latent data.

In problems with missing or latent data, a standard approach is to first impute the unobserved data, then perform all statistical analyses on the completed dataset--corresponding to the observed data and imputed unobserved data--using standard procedures for complete-data inference. Here, we extend this approach to model checking by demonstrating the advantages of the use of completed-data model diagnostics on imputed completed datasets. The approach is set in the theoretical framework of Bayesian posterior predictive checks (but, as with missing-data imputation, our methods of missing-data model checking can also be interpreted as "predictive inference" in a non-Bayesian context). We consider the graphical diagnostics within this framework. Advantages of the completed-data approach include: (1) One can often check model fit in terms of quantities that are of key substantive interest in a natural way, which is not always possible using observed data alone. (2) In problems with missing data, checks may be devised that do not require to model the missingness or inclusion mechanism; the latter is useful for the analysis of ignorable but unknown data collection mechanisms, such as are often assumed in the analysis of sample surveys and observational studies. (3) In many problems with latent data, it is possible to check qualitative features of the model (for example, independence of two variables) that can be naturally formalized with the help of the latent data. We illustrate with several applied examples.

Animals↗

Interpreting small quantities of DNA: the hierarchy of propositions and the use of Bayesian networks.

The dramatic increase in the sensitivity of DNA profiling systems that has occurred over recent years has led to the need to address a wider range of interpretational problems in forensic science. The issues surrounding questions of the kind "whose DNA is this?" have been the subject of considerable controversy but now it is clear that the emphasis is shifting to questions of the kind "how did this DNA get here?" Such issues are discussed in this paper and new insights are provided by two particular recent developments. First, the notion of the "hierarchy of propositions" that has arisen from a project called Case Assessment and Interpretation (CAI) that has been running in the British Forensic Science Service (FSS). Second, a technique for drawing inferences in the face of many interacting considerations, known as "Bayesian networks"--or "Bayes' nets" for short--that has been the subject of an earlier paper in this journal (1). The discussion is carried out by means of case studies, based on actual cases. It is clear that, whereas the inference in relation to the source of the DNA in a crime sample might be overwhelmingly strong, the inference in relation to the propositions that a jury must consider relating to the identity of the actual offender may be much more tentative.

Bayes Theorem↗

Algorithms for inferring haplotypes.

Haplotype phase information in diploid organisms provides valuable information on human evolutionary history and may lead to the development of more efficient strategies to identify genetic variants that increase susceptibility to human diseases. Molecular haplotyping methods are labor-intensive, low-throughput, and very costly. Therefore, algorithms based on formal statistical theories were shown to be very effective and cost-efficient for haplotype reconstruction. This review covers 1) population-based haplotype inference methods: Clark's algorithm, expectation-maximization (EM) algorithm, coalescence-based algorithms (pseudo-Gibbs sampler and perfect/imperfect phylogeny), and partition-ligation algorithm implemented by a fully Bayesian model (Haplotyper) or by EM (PLEM); 2) family-based haplotype inference methods; 3) the handling of genotype scoring uncertainties (i.e., genotyping errors and raw two-dimensional genotype scatterplots) in inferring haplotypes; and 4) haplotype inference methods for pooled DNA samples. The advantages and limitations of each algorithm are discussed. By using simulations based on empirical data on the G6PD gene and TNFRSF5 gene, I demonstrate that different algorithms have different degrees of sensitivity to various extents of population diversities and genotyping error rates. Future development of statistical algorithms for addressing haplotype reconstruction will resort more and more to ideas based on combinatorial mathematics, graphical models, and machine learning, and they will have profound impacts on population genetics and genetic epidemiology with the advent of the human HapMap.

Algorithms↗

A phylogenetic mixture model for detecting pattern-heterogeneity in gene sequence or character-state data.

We describe a general likelihood-based 'mixture model' for inferring phylogenetic trees from gene-sequence or other character-state data. The model accommodates cases in which different sites in the alignment evolve in qualitatively distinct ways, but does not require prior knowledge of these patterns or partitioning of the data. We call this qualitative variability in the pattern of evolution across sites "pattern-heterogeneity" to distinguish it from both a homogenous process of evolution and from one characterized principally by differences in rates of evolution. We present studies to show that the model correctly retrieves the signals of pattern-heterogeneity from simulated gene-sequence data, and we apply the method to protein-coding genes and to a ribosomal 12S data set. The mixture model outperforms conventional partitioning in both these data sets. We implement the mixture model such that it can simultaneously detect rate- and pattern-heterogeneity. The model simplifies to a homogeneous model or a rate-variability model as special cases, and therefore always performs at least as well as these two approaches, and often considerably improves upon them. We make the model available within a Bayesian Markov-chain Monte Carlo framework for phylogenetic inference, as an easy-to-use computer program.

Base Sequence↗

A bivariate quantitative genetic model for a linear Gaussian trait and a survival trait.

With the increasing use of survival models in animal breeding to address the genetic aspects of mainly longevity of livestock but also disease traits, the need for methods to infer genetic correlations and to do multivariate evaluations of survival traits and other types of traits has become increasingly important. In this study we derived and implemented a bivariate quantitative genetic model for a linear Gaussian and a survival trait that are genetically and environmentally correlated. For the survival trait, we considered the Weibull log-normal animal frailty model. A Bayesian approach using Gibbs sampling was adopted. Model parameters were inferred from their marginal posterior distributions. The required fully conditional posterior distributions were derived and issues on implementation are discussed. The two Weibull baseline parameters were updated jointly using a Metropolis-Hasting step. The remaining model parameters with non-normalized fully conditional distributions were updated univariately using adaptive rejection sampling. Simulation results showed that the estimated marginal posterior distributions covered well and placed high density to the true parameter values used in the simulation of data. In conclusion, the proposed method allows inferring additive genetic and environmental correlations, and doing multivariate genetic evaluation of a linear Gaussian trait and a survival trait.

Animals↗

Computational modeling of the Plasmodium falciparum interactome reveals protein function on a genome-wide scale.

Many thousands of proteins encoded by the genome of Plasmodium falciparum, the causal organism of the deadliest form of human malaria, are of unknown function. It is of utmost importance that these proteins be characterized if we are to develop combative strategies against malaria based on the biology of the parasite. In an attempt to infer protein function on a genome-wide scale, we computationally modeled the P. falciparum interactome, elucidating local and global functional relationships between gene products. The resulting interaction network, reconstructed by integrating in silico and experimental functional genomics data within a Bayesian framework, covers approximately 68% of the parasite genome and provides functional inferences for more than 2000 uncharacterized proteins, based on their associations. Network reconstruction involved the use of a novel strategy, where we incorporated continuously updated, uniform reference priors in our Bayesian model. This method for generating interaction maps is thus also well suited for application to other genomes, where pre-existing interactome knowledge is sparse. Additionally, we superimposed this map on genomes of three apicomplexan pathogens--Plasmodium yoelii, Toxoplasma gondii, and Cryptosporidium parvum--describing relationships between these organisms based on retained functional linkages. This comparison provided a glimpse of the highly evolved nature of P. falciparum; for instance, a deficit of nearly 26% in terms of predicted interactions is observed against P. yoelii, because of missing ortholog partners in pairs of functionally linked proteins.

Animals↗

A Bayesian approach to modelling the natural history of a chronic condition from observations with intervention.

To assess the costs and benefits of screening and treatment strategies, it is important to know what would have happened had there been no intervention. In today's ethical climate, however, it is almost impossible to observe this directly and therefore must be inferred from observations with intervention. In this paper, we illustrate a Bayesian approach to this situation when the observations are at separated and unequally spaced time points and the time of intervention is interval censored. We develop a discrete-time Markov model which combines a non-homogeneous Markov chain, used to model the natural progression, with mechanisms that describe the possibility of both treatment intervention and death. We apply this approach to a subpopulation of the Wisconsin Epidemiologic Study of Diabetic Retinopathy, a population-based cohort study to investigate prevalence, incidence, and progression of diabetic retinopathy. In addition, posterior predictive distributions are discussed as a prognostic tool to assist researchers in evaluating costs and benefits of treatment protocols. While we focus this approach on diabetic retinopathy cohort data, we believe this methodology can have wide application.

Age of Onset↗

Bayesian monitoring of phase II trials in cancer chemoprevention.

Early randomized Phase II cancer chemoprevention trials which assess short-term biological activity are critical to the decision process to advance to late Phase II/Phase III trials. We have adapted published Bayesian interim analysis methods (Spiegelhalter et al., J. R. Statist. Soc A, 1994; 157: 357-416) which give greater flexibility and simplicity of inference to the monitoring of randomized controlled Phase II trials using intermediate endpoints. The Bayesian stopping rule is designed to stop the trial more quickly when the evidence suggests ineffectiveness rather than when it suggests biological activity, thus allowing resources to be concentrated on those agents that show the most promise in this early stage of testing. We investigate frequentist performance characteristics of the proposed method through simulation of randomized placebo controlled trials with a growth factor intermediate end-point using mean and variance values derived from the literature. Simulation results show expected error rates and trial size similar to other commonly used group sequential methods for this setting. These results suggest that the Bayesian approach to interim analysis is well suited for monitoring small randomized controlled Phase II chemoprevention trials for early detection of either inactive or promising agents.

Antineoplastic Agents↗

Inferring figure-ground using a recurrent integrate-and-fire neural circuit.

Several theories of early visual perception hypothesize neural circuits that are responsible for assigning ownership of an object's occluding contour to a region which represents the "figure." Previously, we have presented a Bayesian network model which integrates multiple cues and uses belief propagation to infer local figure-ground relationships along an object's occluding contour. In this paper, we use a linear integrate-and-fire model to demonstrate how such inference mechanisms could be carried out in a biologically realistic neural circuit. The circuit maps the membrane potentials of individual neurons to log probabilities and uses recurrent connections to represent transition probabilities. The network's "perception" of figure-ground is demonstrated for several examples, including perceptually ambiguous figures, and compared qualitatively and quantitatively with human psychophysics.

Action Potentials↗

A semi-parametric Bayesian approach to generalized linear mixed models.

The linear mixed effects model with normal errors is a popular model for the analysis of repeated measures and longitudinal data. The generalized linear model is useful for data that have non-normal errors but where the errors are uncorrelated. A descendant of these two models generates a model for correlated data with non-normal errors, called the generalized linear mixed model (GLMM). Frequentist attempts to fit these models generally rely on approximate results and inference relies on asymptotic assumptions. Recent advances in computing technology have made Bayesian approaches to this class of models computationally feasible. Markov chain Monte Carlo methods can be used to obtain 'exact' inference for these models, as demonstrated by Zeger and Karim. In the linear or generalized linear mixed model, the random effects are typically taken to have a fully parametric distribution, such as the normal distribution. In this paper, we extend the GLMM by allowing the random effects to have a non-parametric prior distribution. We do this using a Dirichlet process prior for the general distribution of the random effects. The approach easily extends to more general population models. We perform computations for the models using the Gibbs sampler.

Bayes Theorem↗

A probabilistic approach to large-scale association scans: a semi-Bayesian method to detect disease-predisposing alleles.

Recent analytic and technological breakthroughs have set the stage for genome-wide linkage disequilibrium studies to map disease-susceptibility variants. This paper discusses a probabilistic methodology for making disease-mapping inferences in large-scale case-control genetic studies. The semi-Bayesian approach promoted compares the probability of the observed data under disease hypotheses to the probability of the data under a null hypothesis defined by data at all the markers interrogated in a large study. This method automatically adjusts for the effects of diffuse population stratification. It is claimed that this characterization of the evidence for or against disease models may facilitate more appropriate inductions for large-scale genetic studies. Results include (i) an analytic solution for the population stratification-adjusted Bayes' factor, (ii) the relationship between sample size and Bayes' factors, (iii) an extension to an approximate Bayes' factor calculated across closely-linked sites, and (iv) an extension across multiple studies. Although this paper deals exclusively with genetic studies, it is possible to generalize the approach to treat many different large-scale experiments including studies of gene expression and proteomics.

Journal Article↗

Is there a star tree paradox?

Concerns have been raised that posterior probabilities on phylogenetic trees can be unreliable when the true tree is unresolved or has very short internal branches, because existing methods for Bayesian phylogenetic analysis do not explicitly evaluate unresolved trees. Two recent papers have proposed that evaluating only resolved trees results in a "star tree paradox": when the true tree is unresolved or close to it, posterior probabilities were predicted to become increasingly unpredictable as sequence length grows, resulting in inflated confidence in one resolved tree or another and an increasing risk of false-positive inferences. Here we show that this is not the case; existing Bayesian methods do not lead to an inflation of statistical confidence, provided the evolutionary model is correct and uninformative priors are assumed. Posterior probabilities do not become increasingly unpredictable with increasing sequence length, and they exhibit conservative type I error rates, leading to a low rate of false-positive inferences. With infinite data, posterior probabilities give equal support for all resolved trees, and the rate of false inferences falls to zero. We conclude that there is no star tree paradox caused by not sampling unresolved trees.

Bayes Theorem↗