Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Bayesian analysis and inference from QSAR predictive model results.

QSAR models have been under development for decades but acceptance and utilization of model results have been slow, in part, because there is no widely accepted metric for assessing their reliability. We reapply a method commonly used in quantitative epidemiology and medical decision-making for evaluating the results of screening tests to assess reliability of a QSAR model. It quantifies the accuracy (expressed as sensitivity and specificity) of QSAR models as conditional probabilities of correct and incorrect classification of chemical characteristic, given a true characteristic. Using Bayes formula, these conditional probabilities are combined with prior information to generate a posterior distribution to determine the probability a specific chemical has a particular characteristic, given a model prediction. As an example, we apply this approach to evaluate the predictive reliability of a CATABOL model and base on it a "ready" and "not ready" biodegradability classification. Finally, we show how predictive capability of the model can be improved by sequential use of two models, the first one with high sensitivity and the second with high specificity.

Bayes Theorem↗

Gene networks inference using dynamic Bayesian networks.

This article deals with the identification of gene regulatory networks from experimental data using a statistical machine learning approach. A stochastic model of gene interactions capable of handling missing variables is proposed. It can be described as a dynamic Bayesian network particularly well suited to tackle the stochastic nature of gene regulation and gene expression measurement. Parameters of the model are learned through a penalized likelihood maximization implemented through an extended version of EM algorithm. Our approach is tested against experimental data relative to the S.O.S. DNA Repair network of the Escherichia coli bacterium. It appears to be able to extract the main regulations between the genes involved in this network. An added missing variable is found to model the main protein of the network. Good prediction abilities on unlearned data are observed. These first results are very promising: they show the power of the learning algorithm and the ability of the model to capture gene interactions.

Algorithms↗

Bayesian predictive approach for inference about proportions.

This paper investigates the Bayesian procedures for comparing proportions. These procedures are especially suitable for accepting (or rejecting) the equivalence of two population proportions. Furthermore the Bayesian predictive probabilities provide a natural and flexible tool in monitoring trials, especially for choosing a sample size and for conducting interim analyses. These methods are illustrated with two examples where antithrombotic treatments are administrated to prevent further occurrences of thromboses.

Bayes Theorem↗

Statistical and Bayesian approaches to RNA secondary structure prediction.

Prediction of RNA secondary structure is a fundamental problem in computational structural biology. For several decades, free energy minimization has been the most popular method for prediction from a single sequence. In recent years, the McCaskill algorithm for computation of partition function and base-pair probabilities has become increasingly appreciated. This paradigm-shifting work has inspired the developments of extended partition function algorithms, statistical sampling and clustering, and application of Bayesian statistical inference. The performance of thermodynamics-based methods is limited by thermodynamic rules and parameters. However, further improvements may come from statistical estimates derived from structural databases for thermodynamics parameters with weak or little experimental data. The Bayesian inference approach appears to be promising in this context.

Algorithms↗

Comparing bootstrap and posterior probability values in the four-taxon case.

Assessment of the reliability of a given phylogenetic hypothesis is an important step in phylogenetic analysis. Historically, the nonparametric bootstrap procedure has been the most frequently used method for assessing the support for specific phylogenetic relationships. The recent employment of Bayesian methods for phylogenetic inference problems has resulted in clade support being expressed in terms of posterior probabilities. We used simulated data and the four-taxon case to explore the relationship between nonparametric bootstrap values (as inferred by maximum likelihood) and posterior probabilities (as inferred by Bayesian analysis). The results suggest a complex association between the two measures. Three general regions of tree space can be identified: (1) the neutral zone, where differences between mean bootstrap and mean posterior probability values are not significant, (2) near the two-branch corner, and (3) deep in the two-branch corner. In the last two regions, significant differences occur between mean bootstrap and mean posterior probability values. Whether bootstrap or posterior probability values are higher depends on the data in support of alternative topologies. Examination of star topologies revealed that both bootstrap and posterior probability values differ significantly from theoretical expectations; in particular, there are more posterior probability values in the range 0.85-1 than expected by theory. Therefore, our results corroborate the findings of others that posterior probability values are excessively high. Our results also suggest that extrapolations from single topology branch-length studies are unlikely to provide any general conclusions regarding the relationship between bootstrap and posterior probability values.

Computer Simulation↗

Towards unbiased parentage assignment: combining genetic, behavioural and spatial data in a Bayesian framework.

Inferring the parentage of a sample of individuals is often a prerequisite for many types of analysis in molecular ecology, evolutionary biology and quantitative genetics. In all but a few cases, the method of parentage assignment is divorced from the methods used to estimate the parameters of primary interest, such as mate choice or heritability. Here we present a Bayesian approach that simultaneously estimates the parentage of a sample of individuals and a wide range of population-level parameters in which we are interested. We show that joint estimation of parentage and population-level parameters increases the power of parentage assignment, reduces bias in parameter estimation, and accurately evaluates uncertainty in both. We illustrate the method by analysing a number of simulated test data sets, and through a re-analysis of parentage in the Seychelles warbler, Acrocephalus sechellensis. A combination of behavioural, spatial and genetic data are used in the analyses and, importantly, the method does not require strong prior information about the relationship between nongenetic data and parentage.

Animals↗

Phylogeographic analysis of the cornsnake (Elaphe guttata) complex as inferred from maximum likelihood and Bayesian analyses.

Most phylogeographic studies have used maximum likelihood or maximum parsimony to infer phylogeny and bootstrap analysis to evaluate support for trees. Recently, Bayesian methods using Marlov chain Monte Carlo to search tree space and simultaneously estimate tree support have become popular due to its fast search speed and ability to create a posterior distribution of parameters of interest. Here, I present a study that utilizes Bayesian methods to infer phylogenetic relationships of the cornsnake (Elaphe guttata) complex using cytochrome b sequences. Examination of the posterior probability distributions confirms the existence of three geographic lineages. Additionally, there is no support for the monophyly of the subspecies of E. guttata. Results suggest the three geographic lineages partially conform to the ranges of previously defined subspecies, although Shimodaira-Hasegawa tests suggest that subspecies-constrained trees produce significantly poorer likelihood estimates than the most likely trees reflecting the evolution of three geographic assemblages. Based on molecular support, these three geographic assemblages are recognized as species using evolutionary species criteria: E. guttata, Elaphe slowinskii, and Elaphe emoryi [phylogeographic, maximum likelihood, maximum parsimony, bootstrap, Bayesian, Markov chain Monte Carlo, cornsnake, Cytochrome b, geographic lineages, E. guttta, E. slowinskii, and E. emoryi].

Animals↗

Inferring pathways and networks with a Bayesian framework.

Numerous mathematical methods have been adapted and developed to quantitatively reverse engineer biological networks, for example, signal transduction pathways, from experimental micro-array data. Compared with stochastic methods, such as Boolean networks, and deterministic methods, such as thermodynamic or differential equation-based models, Bayesian network analysis has the ability to assess, with scoring metrics, causal relations based on conditional probabilities and thus permit hypothesis testing. The goal of this paper is to illustrate the integration of several Bayesian based techniques into a unified Bayesian framework that can infer hepatocellular networks from metabolic data. Reverse engineering of pathways and networks provides a framework for predictive modeling and hypotheses testing to gain deeper insight into living organisms, disease mechanisms, and targeted therapeutics. Evaluating this methodology initially against the known biochemical network provides confidence in the networks that are uncovered from the experimental data using this framework. From the metabolic data we inferred the known sub-networks, such as the tricarboxylic acid (TCA) and urea cycles. In addition, we combined the relationships learned from the data and our current knowledge of the biological system to postulate several alternative metabolic sub-network models that can predict a particular cellular function, such as intracellular triglyceride accumulation.

Algorithms↗

Hierarchical Bayesian model for prevalence inferences and determination of a country's status for an animal pathogen.

Certification that a country, region or state is "free" from a pathogen or has a prevalence less than a threshold value has implications for trade in animals and animal products. We develop a Bayesian model for assessment of (i) the probability that a country is "free" of or has an animal pathogen, (ii) the proportion of infected herds in an infected country, and (iii) the within-herd prevalence in infected herds. The model uses test results from animals sampled in a two-stage cluster sample of herds within a country. Model parameters are estimated using modern Markov-chain Monte Carlo methods. We demonstrate our approach using published data from surveys of Newcastle disease and porcine reproductive and respiratory syndrome in Switzerland, and for three simulated data sets.

Animals↗

Approximate Bayesian computation in population genetics.

We propose a new method for approximate Bayesian statistical inference on the basis of summary statistics. The method is suited to complex problems that arise in population genetics, extending ideas developed in this setting by earlier authors. Properties of the posterior distribution of a parameter, such as its mean or density curve, are approximated without explicit likelihood calculations. This is achieved by fitting a local-linear regression of simulated parameter values on simulated summary statistics, and then substituting the observed summary statistics into the regression equation. The method combines many of the advantages of Bayesian statistical inference with the computational efficiency of methods based on summary statistics. A key advantage of the method is that the nuisance parameters are automatically integrated out in the simulation step, so that the large numbers of nuisance parameters that arise in population genetics problems can be handled without difficulty. Simulation results indicate computational and statistical efficiency that compares favorably with those of alternative methods previously proposed in the literature. We also compare the relative efficiency of inferences obtained using methods based on summary statistics with those obtained directly from the data using MCMC.

Bayes Theorem↗

A Bayesian model comparison approach to inferring positive selection.

A popular approach to detecting positive selection is to estimate the parameters of a probabilistic model of codon evolution and perform inference based on its maximum likelihood parameter values. This approach has been evaluated intensively in a number of simulation studies and found to be robust when the available data set is large. However, uncertainties in the estimated parameter values can lead to errors in the inference, especially when the data set is small or there is insufficient divergence between the sequences. We introduce a Bayesian model comparison approach to infer whether the sequence as a whole contains sites at which the rate of nonsynonymous substitution is greater than the rate of synonymous substitution. We incorporated this probabilistic model comparison into a Bayesian approach to site-specific inference of positive selection. Using simulated sequences, we compared this approach to the commonly used empirical Bayes approach and investigated the effect of tree length on the performance of both methods. We found that the Bayesian approach outperforms the empirical Bayes method when the amount of sequence divergence is small and is less prone to false-positive inference when the sequences are saturated, while the results are indistinguishable for intermediate levels of sequence divergence.

Bayes Theorem↗

A coalescence-guided hierarchical Bayesian method for haplotype inference.

Haplotype inference from phase-ambiguous multilocus genotype data is an important task for both disease-gene mapping and studies of human evolution. We report a novel haplotype-inference method based on a coalescence-guided hierarchical Bayes model. In this model, a hierarchical structure is imposed on the prior haplotype frequency distributions to capture the similarities among modern-day haplotypes attributable to their common ancestry. As a consequence, the model both allows distinct haplotypes to have different a priori probabilities according to the inferred hierarchical ancestral structure and results in a proper joint posterior distribution for all the parameters of interest. A Markov chain-Monte Carlo scheme is designed to draw from this posterior distribution. By using coalescence-based simulation and empirically generated data sets (Whitehead Institute's inflammatory bowel disease data sets and HapMap data sets), we demonstrate the merits of the new method in comparison with HAPLOTYPER and PHASE, with or without the presence of recombination hotspots and missing genotypes.

Algorithms↗

Genetic evidence for long-term population decline in a savannah-dwelling primate: inferences from a hierarchical bayesian model.

The purpose of this study was to test for evidence that savannah baboons (Papio cynocephalus) underwent a population expansion in concert with a hypothesized expansion of African human and chimpanzee populations during the late Pleistocene. The rationale is that any type of environmental event sufficient to cause simultaneous population expansions in African humans and chimpanzees would also be expected to affect other codistributed mammals. To test for genetic evidence of population expansion or contraction, we performed a coalescent analysis of multilocus microsatellite data using a hierarchical Bayesian model. Markov chain Monte Carlo (MCMC) simulations were used to estimate the posterior probability density of demographic and genealogical parameters. The model was designed to allow interlocus variation in mutational and demographic parameters, which made it possible to detect aberrant patterns of variation at individual loci that could result from heterogeneity in mutational dynamics or from the effects of selection at linked sites. Results of the MCMC simulations were consistent with zero variance in demographic parameters among loci, but there was evidence for a 10- to 20-fold difference in mutation rate between the most slowly and most rapidly evolving loci. Results of the model provided strong evidence that savannah baboons have undergone a long-term historical decline in population size. The mode of the highest posterior density for the joint distribution of current and ancestral population size indicated a roughly eightfold contraction over the past 1,000 to 250,000 years. These results indicate that savannah baboons apparently did not share a common demographic history with other codistributed primate species.

Animals↗

Fundamental differences between the methods of maximum likelihood and maximum posterior probability in phylogenetics.

Using a four-taxon example under a simple model of evolution, we show that the methods of maximum likelihood and maximum posterior probability (which is a Bayesian method of inference) may not arrive at the same optimal tree topology. Some patterns that are separately uninformative under the maximum likelihood method are separately informative under the Bayesian method. We also show that this difference has impact on the bootstrap frequencies and the posterior probabilities of topologies, which therefore are not necessarily approximately equal. Efron et al. (Proc. Natl. Acad. Sci. USA 93:13429-13434, 1996) stated that bootstrap frequencies can, under certain circumstances, be interpreted as posterior probabilities. This is true only if one includes a non-informative prior distribution of the possible data patterns, and most often the prior distributions are instead specified in terms of topology and branch lengths. [Bayesian inference; maximum likelihood method; Phylogeny; support.].

Bayes Theorem↗

Estimating haplotype frequencies and standard errors for multiple single nucleotide polymorphisms.

Estimating haplotype frequencies becomes increasingly important in the mapping of complex disease genes, as millions of single nucleotide polymorphisms (SNPs) are being identified and genotyped. When genotypes at multiple SNP loci are gathered from unrelated individuals, haplotype frequencies can be accurately estimated using expectation-maximization (EM) algorithms (Excoffier and Slatkin, 1995; Hawley and Kidd, 1995; Long et al., 1995), with standard errors estimated using bootstraps. However, because the number of possible haplotypes increases exponentially with the number of SNPs, handling data with a large number of SNPs poses a computational challenge for the EM methods and for other haplotype inference methods. To solve this problem, Niu and colleagues, in their Bayesian haplotype inference paper (Niu et al., 2002), introduced a computational algorithm called progressive ligation (PL). But their Bayesian method has a limitation on the number of subjects (no more than 100 subjects in the current implementation of the method). In this paper, we propose a new method in which we use the same likelihood formulation as in Excoffier and Slatkin's EM algorithm and apply the estimating equation idea and the PL computational algorithm with some modifications. Our proposed method can handle data sets with large number of SNPs as well as large numbers of subjects. Simultaneously, our method estimates standard errors efficiently, using the sandwich-estimate from the estimating equation, rather than the bootstrap method. Additionally, our method admits missing data and produces valid estimates of parameters and their standard errors under the assumption that the missing genotypes are missing at random in the sense defined by Rubin (1976).

Algorithms↗