Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Inference of demographic history from genealogical trees using reversible jump Markov chain Monte Carlo.

BACKGROUND: Coalescent theory is a general framework to model genetic variation in a population. Specifically, it allows inference about population parameters from sampled DNA sequences. However, most currently employed variants of coalescent theory only consider very simple demographic scenarios of population size changes, such as exponential growth. RESULTS: Here we develop a coalescent approach that allows Bayesian non-parametric estimation of the demographic history using genealogies reconstructed from sampled DNA sequences. In this framework inference and model selection is done using reversible jump Markov chain Monte Carlo (MCMC). This method is computationally efficient and overcomes the limitations of related non-parametric approaches such as the skyline plot. We validate the approach using simulated data. Subsequently, we reanalyze HIV-1 sequence data from Central Africa and Hepatitis C virus (HCV) data from Egypt. CONCLUSIONS: The new method provides a Bayesian procedure for non-parametric estimation of the demographic history. By construction it additionally provides confidence limits and may be used jointly with other MCMC-based coalescent approaches.

Algorithms↗

Innovations in bayes and empirical bayes methods: estimating parameters, populations and ranks.

By formalizing the relation among components and 'borrowing information' among them, Bayes and empirical Bayes methods can produce more valid, efficient and informative statistical evaluations than those based on traditional methods. In addition, Bayesian structuring of complicated models and goals guides development of appropriate statistical approaches and generates summaries which properly account for sampling and modelling uncertainty. Computing innovations enable implementation of complex and relevant models, thereby substantially increasing the role of Bayes/empirical Bayes methods in important statistical assessments. Policy-relevant statistical assessments involve synthesis of information from a set of related components such as medical clinics, geographic regions or research studies. Typical assessments include inference for individual parameters, synthesis over the collection of components (for example, the parameter histogram) and comparisons among parameters (for example, ranks). The relative importance of these goals depends on the context. Bayesian structuring provides a guide to valid inference. For example, while posterior means are the 'obvious' and optimal estimates for individual components under squared error loss, their empirical distribution function (EDF) is underdispersed and never valid for estimating the EDF of the true, underlying parameters. Effective histogram estimates result from optimizing a loss function based in a distance between the histogram and its estimate. Similarly, ranking observed data usually produces poor estimates and ranking posterior means can be inappropriate. Effective estimates should be based on a loss function that caters directly to ranks. Using examples of 'borrowing information', shrinkage and the variance/bias trade-off we motivate Bayes and empirical Bayes analysis. Then, we outline the formal approach and discuss 'triple-goal' estimates with values that when ranked produce optimal ranks, for which the EDF is an optimal estimate of the parameter EDF and such that the values themselves are effective estimates of co-ordinate-specific parameters. We use basic models and data analysis examples to highlight the conceptual and structural issues.

Animals↗

Implementing Gaussian process inference with neural networks.

Gaussian processes compare favourably with backpropagation neural networks as a tool for regression, and Bayesian neural networks have Gaussian process behaviour when the number of hidden neurons tends to infinity. We describe a simple recurrent neural network with connection weights trained by one-shot Hebbian learning. This network amounts to a dynamical system which relaxes to a stable state in which it generates predictions identical to those of Gaussian process regression. In effect an infinite number of hidden units in a feed-forward architecture can be replaced by a merely finite number, together with recurrent connections.

Algorithms↗

Estimation of fugitive lead emission rates from secondary lead facilities using hierarchical Bayesian models.

Fugitive emissions from secondary lead recovery facilities are difficult to estimate and can vary significantly from site to site. A methodology is presented for estimating fugitive emissions using back inference from observed ambient concentrations at nearby monitors, in conjunction with an atmospheric transport and dispersion model. Observed concentrations are regressed against unit source-monitor transfer terms computed by the model, and the fitted parameters of the regression equation include the background ambient lead concentration, the fugitive lead emission rate, and (when stack emissions are assumed to be unknown) the stack lead emission rate. The methodology is implemented at three sites, one each in Florida, Texas, and New York. A hierarchical Bayesian method is used to estimate the parameters of the model, allowing inferences to be made for both site-specific values and multisite (national) distributions of fugitive emissions and background concentrations. Informed prior distributions must be specified for the background lead concentrations and for fugitive and stack emission rates in order to obtain stable estimates. Sensitivity analyses with alternative priors indicate that posterior estimates of background concentrations and fugitive emission rates are relatively insensitive to the assumed priors, although estimated stack emission rates can vary with alternative priors, especially for the New York facility, where the stack emission rate is highly uncertain and poorly resolved by the model. The fugitive lead emission rates estimated for the sites are comparable to, or in some cases (especially Texas and New York) likely larger than the stack emissions that are determined for these facilities. An aggregate predictive distribution is derived for the average fugitive lead emission rate from secondary lead smelting facilities, with a median value of 9.2 x 10(-7) g Pb/m2/sec, and a 90% credible interval from 2.1 x 10(-7)-5.3 x 10(-6) g Pb/m2/sec. This wide range reflects both the variation in fugitive lead emissions from site to site and the high degree of uncertainty resulting from an estimate based on only a very small sample of sites. As such, the primary contribution of this study is methodological, demonstrating how information from multiple sites can be combined and considered simultaneously for the estimation of fugitive emission rates, but recognizing that additional sites must be included to obtain a more precise characterization.

Air Pollutants↗

Inferring evolutionary signals from ecological data in a plant-pathogen metapopulation.

We followed the dynamics of local epidemics in three populations of a natural plant-pathogen system for four sequential years. We characterize the overwintering process with spatial statistics and use a stochastic, spatially explicit, modeling approach with Bayesian parameter estimation to study the spread of the infection during the growing season. Our modeling approach allows us to infer coevolutionary signals from spatiotemporal data on pathogen prevalence. Most importantly, we are able to assess the distribution of resistant hosts within the distribution of all host plants. We show that resistant hosts occur in areas with high pathogen encounter rates, and that the occurrence of resistance correlates with overwintering probability of the pathogen. The estimates for essentially all model parameters are characterized by a large amount of variation over the years and the populations. While the variation in the fraction of resistant hosts and in the force of infection is to a large extent explained by the population, the other model parameters (two parameters describing the shape of the dispersal kernel) vary essentially in an unpredictable manner, suggesting that much of the variation may occur at very fine spatial and temporal scales.

Bayes Theorem↗

Novel conopeptides of the I-superfamily occur in several clades of cone snails.

The I-superfamily of conotoxins represents a new class of peptides in the venom of some Conus species. These toxins are characterized by four disulfide bridges and inhibit or modify ion channels of nerve cells. When testing venoms from 11 Conus species for a functional characterization, blocking activity on potassium channels (like Kv1.1 and Kv1.3 channels, but not Kv1.2 channels) was detected in the venom of Conus capitaneus, Conus miles, Conus vexillum and Conus virgo. Analysis at the cDNA level of these venoms using primers designed according to the amino acid sequence of a potassium channel blocking toxin (ViTx) from C. virgo confirmed the presence of structurally homologous peptides in these venoms. Moreover, peptides belonging to the I-superfamily, but with divergent amino acid sequences, were found in Conus striatus and Conus imperialis. In all cases, the sequences of the precursors' prepro-regions exhibited high conservation, whereas the sequences of the mature peptides ranged from almost identical to highly divergent between species. We then performed phylogenetic analyses of new and published mitochondrial 16S rDNA sequences representing 104 haplotypes from these and numerous other Conus species, using Bayesian, maximum-likelihood, maximum-parsimony and neighbor-joining methods of inference. Cone snails known to possess I-superfamily toxins were assigned to five different major clades in all of the resulting gene trees. Moreover, I-superfamily conopeptides were detected both in vermivorous and piscivorous species of Conus, thus demonstrating the widespread presence of such toxins in this speciose genus beyond evolutionary and ecological groups.

Animals↗

Phylogenetic relationships of the genus Phanerochaete inferred from the internal transcribed spacer region.

Phanerochaete is a genus of resupinate homobasidiomycetes that are saprophytic on woody debris and logs. Morphological studies in the past indicated that Phanerochaete is a heterogeneous assemblage of species. In this study the internal transcribed spacer (ITS) region of the nuclear ribosomal DNA was used to test the monophyly of the genus Phanerocthaete and to infer phylogenetic relationships of the 24 taxa studied. Maximum parsimony, maximum likelihood, and Bayesian analyses do not support the monophyly of the genus. However, a core group of species represented by Phanerochaete velutina, P. chrysosporium, P. sordida, P. sanguinea and others are closely related and group together in a clade. Other common Phanerochaete species including Phanerochaete rimosa, P. chrysorhiza, P. omnivora, P. avellanea, P. tiberculata, P. flava, and P. allantospora, however, do not cluster with the core Phanerochaete group.

DNA, Fungal↗

Evolution of homologous recombination rates across bacteria.

Bacteria are nonsexual organisms but are capable of exchanging DNA at diverse degrees through homologous recombination. Intriguingly, the rates of recombination vary immensely across lineages where some species have been described as purely clonal and others as "quasi-sexual." However, estimating recombination rates has proven a difficult endeavor and estimates often vary substantially across studies. It is unclear whether these variations reflect natural variations across populations or are due to differences in methodologies. Consequently, the impact of recombination on bacterial evolution has not been extensively evaluated and the evolution of recombination rate-as a trait-remains to be accurately described. Here, we developed an approach based on Approximate Bayesian Computation that integrates multiple signals of recombination to estimate recombination rates. We inferred the rate of recombination of 162 bacterial species and one archaeon and tested the robustness of our approach. Our results confirm that recombination rates vary drastically across bacteria; however, we found that recombination rate-as a trait-is conserved in several lineages but evolves rapidly in others. Although some traits are thought to be associated with recombination rate (e.g., GC-content), we found no clear association between genomic or phenotypic traits and recombination rate. Overall, our results provide an overview of recombination rate, its evolution, and its impact on bacterial evolution.

Bacteria↗

Bayesian analysis of response to selection: a case study using litter size in Danish Yorkshire pigs.

Implementation of a Bayesian analysis of a selection experiment is illustrated using litter size [total number of piglets born (TNB)] in Danish Yorkshire pigs. Other traits studied include average litter weight at birth (WTAB) and proportion of piglets born dead (PRBD). Response to selection for TNB was analyzed with a number of models, which differed in their level of hierarchy, in their prior distributions, and in the parametric form of the likelihoods. A model assessment study favored a particular form of an additive genetic model. With this model, the Monte Carlo estimate of the 95% probability interval of response to selection was (0.23; 0.60), with a posterior mean of 0.43 piglets. WTAB showed a correlated response of -7.2 g, with a 95% probability interval equal to (-33.1; 18.9). The posterior mean of the genetic correlation between TNB and WTAB was -0.23 with a 95% probability interval equal to (-0.46; -0.01). PRBD was studied informally; it increases with larger litters, when litter size is >7 piglets born. A number of methodological issues related to the Bayesian model assessment study are discussed, as well as the genetic consequences of inferring response to selection using additive genetic models.

Animal Husbandry↗

Identifiability, exchangeability, and epidemiological confounding.

Non-identifiability of parameters is a well-recognized problem in classical statistics, and Bayesian statisticians have long recognized the importance of exchangeability assumptions in making statistical inferences. A seemingly unrelated problem in epidemiology is that of confounding: bias in estimation of the effects of an exposure on disease risk, due to inherent differences in risk between exposed and unexposed individuals. Using a simple deterministic model for exposure effects, a logical connection is drawn between the concepts of identifiability, exchangeability, and confounding. This connection allows one to view the problem of confounding as arising from problems of identifiability, and reveals the exchangeability assumptions that are implicit in confounder control methods. It also provides further justification for confounder definitions based on comparability of exposure groups, as opposed to collapsibility-based definitions.

Bayes Theorem↗

Prior convictions: Bayesian approaches to the analysis and interpretation of clinical megatrials.

Large, randomized clinical trials ("megatrials") are key drivers of modern cardiovascular practice, since they are cited frequently as the authoritative foundation for evidence-based management policies. Nevertheless, fundamental limitations in the conventional approach to statistical hypothesis testing undermine the scientific basis of the conclusions drawn from these trials. This review describes the conventional approach to statistical inference, highlights its limitations, and proposes an alternative approach based on Bayes' theorem. Despite its inherent subjectivity, the Bayesian approach possesses a number of practical advantages over the conventional approach: 1). it allows the explicit integration of previous knowledge with new empirical data; 2). it avoids the inevitable misinterpretations of p values derived from megatrial populations; and 3). it replaces the misleading p value with a summary statistic having a natural, clinically relevant interpretation-the probability that the study hypothesis is true given the observations. This posterior probability thereby quantifies the likelihood of various magnitudes of therapeutic benefit rather than the single null magnitude to which the p value refers, and it lends itself to graphical sensitivity analyses with respect to its underlying assumptions. Accordingly, the Bayesian approach should be employed more widely in the design, analysis, and interpretation of clinical megatrials.

Bayes Theorem↗

Bayesian cost-effectiveness analysis from clinical trial data.

A key tool for assessing the relative cost-effectiveness of two treatments in health economics is the incremental C/E acceptability curve. We present Bayesian computations for this curve in the case where data on both costs and efficacy are available from a clinical trial. Analysis is given under various formulations of prior information. A case study is analysed in which reasonable prior information is shown to strengthen substantially the posterior inference, leading to a more conclusive assessment of cost-effectiveness. Calculations can be performed using readily available Bayesian software.

Anti-Asthmatic Agents↗

Implications of empirical Bayes meta-analysis for test validation.

Empirical Bayes meta-analysis provides a useful framework for examining test validation. The fixed-effects case in which rho has a single value corresponds to the inference that the situational specificity hypothesis can be rejected in a validity generalization study. A Bayesian analysis of such a case provides a simple and powerful test of rho = 0; such a test has practical implications for significance testing in test validation. The random-effects case in which sigma2rho > 0 provides an explicit method with which to assess the relative importance of local validity studies and previous meta-analyses. Simulated data are used to illustrate both cases. Results of published meta-analyses are used to show that local validation becomes increasingly important as sigma2rho increases. The meaning of the term validity generalization is explored, and the problem of what can be inferred about test transportability in the random-effects case is described.

Bayes Theorem↗

Comparative genome analyses of Arabidopsis spp.: inferring chromosomal rearrangement events in the evolutionary history of A. thaliana.

Comparative genome analysis is a powerful tool that can facilitate the reconstruction of the evolutionary history of the genomes of modern-day species. The model plant Arabidopsis thaliana with its n = 5 genome is thought to be derived from an ancestral n = 8 genome. Pairwise comparative genome analyses of A. thaliana with polyploid and diploid Brassicaceae species have suggested that rapid genome evolution, manifested by chromosomal rearrangements and duplications, characterizes the polyploid, but not the diploid, lineages of this family. In this study, we constructed a low-density genetic linkage map of Arabidopsis lyrata ssp. lyrata (A. l. lyrata; n = 8, diploid), the closest known relative of A. thaliana (MRCA approximately 5 Mya), using A. thaliana-specific markers that resolve into the expected eight linkage groups. We then performed comparative Bayesian analyses using raw mapping data from this study and from a Capsella study to infer the number and nature of rearrangements that distinguish the n = 8 genomes of A. l. lyrata and Capsella from the n = 5 genome of A. thaliana. We conclude that there is strong statistical support in favor of the parsimony scenarios of 10 major chromosomal rearrangements separating these n = 8 genomes from A. thaliana. These chromosomal rearrangement events contribute to a rate of chromosomal evolution higher than previously reported in this lineage. We infer that at least seven of these events, common to both sets of data, are responsible for the change in karyotype and underlie genome reduction in A. thaliana.

Arabidopsis↗

Bayesian analysis of the linear reaction norm model with unknown covariates.

The reaction norm model is becoming a popular approach for the analysis of genotype x environment interactions. In a classical reaction norm model, the expression of a genotype in different environments is described as a linear function (a reaction norm) of an environmental gradient or value. An environmental value is typically defined as the mean performance of all genotypes in the environment, which is usually unknown. One approximation is to estimate the mean phenotypic performance in each environment and then treat these estimates as known covariates in the model. However, a more satisfactory alternative is to infer environmental values simultaneously with the other parameters of the model. This study describes a method and its Bayesian Markov Chain Monte Carlo implementation that makes this possible. Frequentist properties of the proposed method are tested in a simulation study. Estimates of parameters of interest agree well with the true values. Further, inferences about genetic parameters from the proposed method are similar to those derived from a reaction norm model using true environmental values. On the other hand, using phenotypic means as proxies for environmental values results in poor inferences.

Bayes Theorem↗

Helping doctors to draw appropriate inferences from the analysis of medical studies.

Most clinicians and many medical statisticians interpret standard frequentist confidence intervals by invoking the Bayesian concept of subjective probability. Fortunately, the assumptions that render this interpretation acceptable are often quite reasonable in the setting of the practical day-to-day analysis of medical data. This article takes the subjective interpretation of confidence intervals to its logical conclusion and argues that the inferential understanding of clinicians and public health physicians could potentially be improved if, where it was appropriate, standard inferential statements--point estimates, 95 per cent confidence intervals and P-values--were supplemented by estimates of the subjective posterior probability, assuming a uniform prior density, that the true value of a parameter to be estimated exceeds one or a series of thresholds that are clinically critical or easily interpretable. Many decision makers in the health care arena draw totally inappropriate inferences from analyses where the point estimate indicates a clinically valuable effect but the null hypothesis cannot formally be rejected, and, although the proposed approach could be of potential value in a range of settings, it is argued that it could be of particular use in the rational interpretation of underpowered studies that must inform critical clinical or public health decisions.

Bayes Theorem↗

Inference of ancestry: constructing hierarchical reference populations and assigning unknown individuals.

The ability to infer personal genetic ancestry is being increasingly utilised in certain medical and forensic situations. Herein, the unsupervised Bayesian clustering algorithms structure, is employed to analyse 377 autosomal short tandem repeats typed on 1,056 individuals from the Centre d'Etude du Polymorphisme Humain Human Diversity Panel. Individuals of known geographical origin were hierarchically classified into a framework of increasingly homogeneous clusters to serve as reference populations into which individuals of unknown ancestry can be assigned. The groupings were characterised by the geographical affinities of cluster members and the accuracy of these procedures was verified using several genetic indices. Fine-scale substructure was detectable beyond the broad population level classifications that previously have been explored in this dataset. Metrics indicated that within certain lines, the strongest structuring signals were detected at the leaves of the hierarchy where lineage-specific groupings were identified. The accuracy of unknown assignment was assessed at each level of the hierarchy using a 'leave one out' strategy in which each individual was stripped of cluster membership and then re-assigned using the supervised Bayesian clustering algorithm implemented in GeneClass2. Although most clusters at all levels of resolution experienced highly accurate assignment, a decline was observed in the finer levels due to the mixed membership characteristics of some individuals. The parameters defined by this study allowed for assignment of unknown individuals to genetically defined clusters with measured likelihood. Shared ancestry data can then be inferred for the unknown individual.

Algorithms↗

Empirical bayes methods and false discovery rates for microarrays.

In a classic two-sample problem, one might use Wilcoxon's statistic to test for a difference between treatment and control subjects. The analogous microarray experiment yields thousands of Wilcoxon statistics, one for each gene on the array, and confronts the statistician with a difficult simultaneous inference situation. We will discuss two inferential approaches to this problem: an empirical Bayes method that requires very little a priori Bayesian modeling, and the frequentist method of "false discovery rates" proposed by Benjamini and Hochberg in 1995. It turns out that the two methods are closely related and can be used together to produce sensible simultaneous inferences.

Bayes Theorem↗