Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Directed Convergence in Stable Percept Acquisition.

We view a perceptual capacity as a nondeductive inference, represented as a function from a set of premises to a set of conclusions. The application of the function to a single premise to produce a single conclusion is called a "percept" or "instantaneous percept." We define a stable percept as a convergent sequence of instantaneous percepts. Assuming that the sets of premises and conclusions are metric spaces, we introduce a strategy for acquiring stable percepts, called directed convergence. We consider probabilistic inferences, where the premise and conclusion sets are spaces of probability measures, and in this context we study Bayesian probabilistic/recursive inference. In this type of Bayesian inference the premises are probability measures, and the prior as well as the posterior is updated nontrivially at each iteration. This type of Bayesian inference is distinguished from classical Bayesian statistical inference where the prior remains fixed, and the posterior evolves by conditioning on successively more punctual premises. We indicate how the directed convergence procedure may be implemented in the context of Bayesian probabilistic/recursive inference. We discuss how the L(infinity) metric can be used to give numerical control of this type of Bayesian directed convergence. Copyright 2001 Academic Press.

Journal Article↗

An introduction to the Bayesian analysis of clinical trials.

Although most clinical trials comparing therapies are analyzed using classical hypothesis testing and P values, such methods do not yield the information most useful to the clinician, that is, the probability that one treatment is more efficacious than another. Bayesian inference can yield this probability but only if we quantify our prior beliefs about the possible efficacies of the treatments studied. This article gives a brief introduction to Bayesian methods and contrasts them with classical hypothesis testing. It shows that the quantification of prior beliefs is a common and necessary part of the interpretation of clinical information, whether from a laboratory test or published clinical trial. Advantages of Bayesian analysis over classical analysis of clinical trials include the ability to incorporate prior information regarding treatment efficacies into the analysis; the ability to make multiple unscheduled inspections of accumulating data without increasing the error rate of the study; and the ability to calculate the probability that one treatment is more effective than another. Because it is likely that Bayesian methods will be used more often in the analysis of future clinical trials, investigators and readers should be aware of the two schools of statistical thought and the strengths and weaknesses of each.

Bayes Theorem↗

Bayesian framework for least-squares support vector machine classifiers, gaussian processes, and kernel Fisher discriminant analysis.

The Bayesian evidence framework has been successfully applied to the design of multilayer perceptrons (MLPs) in the work of MacKay. Nevertheless, the training of MLPs suffers from drawbacks like the nonconvex optimization problem and the choice of the number of hidden units. In support vector machines (SVMs) for classification, as introduced by Vapnik, a nonlinear decision boundary is obtained by mapping the input vector first in a nonlinear way to a high-dimensional kernel-induced feature space in which a linear large margin classifier is constructed. Practical expressions are formulated in the dual space in terms of the related kernel function, and the solution follows from a (convex) quadratic programming (QP) problem. In least-squares SVMs (LS-SVMs), the SVM problem formulation is modified by introducing a least-squares cost function and equality instead of inequality constraints, and the solution follows from a linear system in the dual space. Implicitly, the least-squares formulation corresponds to a regression formulation and is also related to kernel Fisher discriminant analysis. The least-squares regression formulation has advantages for deriving analytic expressions in a Bayesian evidence framework, in contrast to the classification formulations used, for example, in gaussian processes (GPs). The LS-SVM formulation has clear primal-dual interpretations, and without the bias term, one explicitly constructs a model that yields the same expressions as have been obtained with GPs for regression. In this article, the Bayesian evidence framework is combined with the LS-SVM classifier formulation. Starting from the feature space formulation, analytic expressions are obtained in the dual space on the different levels of Bayesian inference, while posterior class probabilities are obtained by marginalizing over the model parameters. Empirical results obtained on 10 public domain data sets show that the LS-SVM classifier designed within the Bayesian evidence framework consistently yields good generalization performances.

Artificial Intelligence↗

Bayesian estimation of range for microsatellite loci.

Microsatellite loci have become important in population genetics because of their high level of polymorphism in natural populations, very frequent occurrence throughout the genome, and apparently high mutation rate. Observed repeat numbers (alleles size) in natural populations and expectations based on computer simulations suggest that the range of repeat numbers at a microsatellite locus is restricted. This range is a key parameter that should be properly estimated in order to proceed with calculations of divergence times in phylogenetic studies and to better investigate the within- and between-population variability. The 'plug-in' estimate of range based on the minimum and maximum value observed in a sample is not satisfactory because of the relatively large number of alleles in comparison with typical sample sizes. In this paper, a set of data from 30 dinucleotide microsatellite loci is analysed under the assumption of independence among loci. Bayesian inference on range for one locus is obtained by assuming that constraints on range values exist as sharp bounds. Closed-form calculations and robustness revealed by our analysis suggest that the proposed Bayesian approach might be routinely used by researchers to classify microsatellite loci according to the estimated value of their allelic range.

Bayes Theorem↗

The Bayesian approach improves the electrocardiographic diagnosis of broad complex tachycardia.

Despite numerous attempts at devising algorithms for diagnosing broad complex tachycardia (BCT) on the basis of the electrocardiogram (ECG), misdiagnosis is still common. The reason for this may lie with difficulty in implementing existent algorithms in practice, due to imperfect ascertainment of ECG features within them. An attempt was made to approach the problem afresh with the Bayesian inference by the construction of a diagnostic algorithm centered around the likelihood ratio (LR). Previously studied ECG features most effective in discriminating ventricular tachycardia (VT) from supraventricular tachycardia with aberrant conduction (SVTAC), according to their LR values, were selected for inclusion into a Bayesian diagnostic algorithm. A test set of 244 BCT ECGs was assembled and shown to three independent observers who were blinded to the diagnoses made at electrophysiological study. Their diagnostic accuracy by the Bayesian algorithm was compared against that by clinical judgement with the diagnoses from EPS as the criterial standard. Clinical judgement correctly diagnosed 35% of SVTAC, 85% of VT, and 47% of fascicular tachycardia. In comparison, by the Bayesian algorithm devised, 52% of SVTAC, 95% of VT, and 97% of fascicular tachycardia were correctly diagnosed. The Bayesian algorithm devised has proved to be superior to the clinical judgement of the observers who participated in this study, and theoretically will obviate the problem of imperfect ascertainment of ECG features. Hence, it holds the promise for being an effective tool for routine use in clinical practice.

Algorithms↗

The diagnosis of polyarteritis nodosa. I. A literature-based decision analysis approach.

We investigated diagnostic testing in polyarteritis nodosa (PAN) by calculating, from published data, the sensitivity and specificity of visceral angiography and muscle, nerve, testicle, kidney, and liver biopsy. Test sequence strategies were constructed by Bayesian inference using a computer program written for this purpose. Test sequences were compared with an aggressive strategy consisting of repeated tests until there was a positive finding or until the available tests were exhausted, and a conservative strategy consisting of 1 biopsy procedure plus angiography. The Bayesian analysis agreed most closely with the conservative approach for most prior probabilities (degree of suspicion) that a patient had PAN. The aggressive strategy had an overall sensitivity of 90% and specificity of 91%, whereas the conservative strategy was 85% sensitive and 96% specific. Furthermore, the aggressive strategy was more costly ($2,986 versus $1,961) and had a higher rate of morbidity (3.8 versus 2.7 days of hospitalization per patient evaluated) than did the conservative strategy. The mortality rates of both strategies were equivalent (approximately 0.05 deaths per hundred patients evaluated). The per-case cost of diagnosis increased as prevalence decreased, and at 10% prevalence, the aggressive strategy cost more than $17,000 per case diagnosed. Sensitivity analysis revealed that the strategies were moderately affected by the test characteristics, within reasonable assumptions, but that the differences in conservative and aggressive approaches remained. Thus, our analysis based on available data and the assumption of test independence suggests that the preferred diagnostic evaluation of patients with symptoms suggestive of PAN consists, in most cases, of a single biopsy procedure, with angiographic evaluation if necessary.

Costs and Cost Analysis↗

Generalized linear mixed models in dairy cattle breeding.

Fitness and fertility traits of dairy cattle are of increasing importance and are often measured on a discrete scale. The development and application of generalized linear mixed models to the genetic analysis of these traits are reviewed. Because current genetic evaluation systems are predominantly based on animal models, the inferential challenges of highly parameterized generalized linear mixed models are discussed. Development and adoption of new methods for drawing appropriate inferences on dispersion parameters are essential. Recent hierarchical extensions have been proposed for generalized linear mixed models, allowing for complex dispersion patterns that accommodate heteroscedasticity and outlier robustness. Steady advances in available computing power have facilitated multiple-trait analyses involving continuous and discrete measures. Full Bayesian inference via the development of Markov Chain Monte Carlo methods will continue to allow even greater generality and dimensions in the genetic model.

Animals↗

Baseline risk as predictor of treatment benefit: three clinical meta-re-analyses.

A relationship between baseline risk and treatment effect is increasingly investigated as a possible explanation of between-study heterogeneity in clinical trial meta-analysis. An approach that is still often applied in the medical literature is to plot the estimated treatment effects against the estimated measures of risk in the control groups (as a measure of baseline risk), and to compute the ordinary weighted least squares regression line. However, it has been pointed out by several authors that this approach can be seriously flawed. The main problem is that the observed treatment effect and baseline risk measures should be viewed as estimates rather than the true values. In recent years several methods have been proposed in the statistical literature to potentially deal with the measurement errors in the estimates. In this article we propose a vague priors Bayesian solution to the problem which can be carried out using the 'Bayesian inference using Gibbs sampling' (BUGS) implementation of Markov chain Monte Carlo numerical integration techniques. Different from other proposed methods, it uses the exact rather than an approximate likelihood, while it can handle many different treatment effect measures and baseline risk measures. The method differs from a recently proposed Bayesian method in that it explicitly models the distribution of the underlying baseline risks. We apply the method to three meta-analyses published in the medical literature and compare the results with the outcomes of the other recently proposed methods. In particular we compare our approach to McIntosh's method, for which we show how it can be carried out using standard statistical software. We conclude that our proposed method offers a very general and flexible solution to the problem, which can be carried out relatively easily with existing Bayesian analysis software. A confidence band for the underlying relationship between true effect measure and baseline risk and a confidence interval for the value of the baseline risk measure for which there is no treatment effect are easily obtained by-products of our approach.

Bayes Theorem↗

Hypermedia and randomized algorithms for medical expert systems.

KNET is an environment for constructing probabilistic, knowledge-intensive systems within the axiomatic framework of decision theory. The KNET architecture defines a complete separation between the hypermedia user interface on the one hand, and the representation and management of expert opinion on the other. KNET offers a choice of algorithms for probabilistic inference. We and our coworkers have used KNET to build consultation systems for lymph-node pathology, bone-marrow transplantation therapy, clinical epidemiology, and alarm management in the intensive-care unit. Most important, KNET contains a randomized approximation scheme (RAS) for the difficult and almost certainly intractable problem of Bayesian inference. Our algorithm can, in many circumstances, perform efficient approximate inference in large and richly interconnected models of medical diagnosis. In this article, we describe the architecture of KNET, construct a randomized algorithm for probabilistic inference, and analyze the algorithm's performance. Finally, we characterize our algorithms' empiric behavior and explore its potential for parallel speedups. From design to implementation, then, KNET demonstrates the crucial interaction between theoretical computer science and medical informatics.

Algorithms↗

Predicting dose-time profiles of solar energetic particle events using Bayesian forecasting methods.

Bayesian inference techniques, coupled with Markov chain Monte Carlo sampling methods, are used to predict dose-time profiles for energetic solar particle events. Inputs into the predictive methodology are dose and dose-rate measurements obtained early in the event. Surrogate dose values are grouped in hierarchical models to express relationships among similar solar particle events. Models assume nonlinear, sigmoidal growth for dose throughout an event. Markov chain Monte Carlo methods are used to sample from Bayesian posterior predictive distributions for dose and dose rate. Example predictions are provided for the November 8, 2000, and August 12, 1989, solar particle events.

Bayes Theorem↗

MDL and the statistical mechanics of protein potentials.

The combination of a wealth of structural data and impressive computational power provides detailed information pertaining to the structure and dynamics of biomacromolecules. A natural inclination is to incorporate this information into models to gain added predictive power on protein folding and stability. There has been considerable recent interest in developing "knowledge-based" potentials to describe internal interactions in proteins. In these approaches, probability distribution functions are inferred from existing knowledge. A common assumption has been the "quasi-chemical approximation" or "Boltzmann device". This method relates statistical mechanical probabilities to observed frequencies. The validity of this approach is discussed in detail from a statistical mechanics perspective. Because statistical mechanics is a form of statistical inference based on a lack of knowledge of the system, the "Boltzmann device" does not have a rigorous theoretical justification. In the present work, a statistical mechanics based on partial knowledge of the system is employed. This statistical mechanical scheme uses the minimum description length (MDL) of phase space as its main tool. With this approach, "knowledge-based" potentials can be derived in a rigorous fashion. In practical calculations, these potentials are best obtained using Bayesian inference methods similar to those used in image reconstruction.

Algorithms↗

The problem of multiple inference in studies designed to generate hypotheses.

Epidemiologic research often involves the simultaneous assessment of associations between many risk factors and several disease outcomes. In such situations, often designed to generate hypotheses, multiple univariate hypothesis-testing is not an appropriate basis for inference. The number of true positive associations in a collection of many associations can be estimated by comparing the observed distribution of p values for the positive associations to a theoretical uniform distribution, or to the observed distribution of negative associations, or to an empiric randomization distribution. None of these approaches, however, will distinguish the true from the false positive associations. Various criteria for selecting a subset of associations to report are considered by the authors, including Bonferoni adjustment of p values, splitting the sample for searching and testing, Bayesian inference, and decision theory. The authors prefer an approach in which all associations in the data are reported, whether significant or not, followed by a ranking in order of priority for investigation using empirical Bayes techniques. Methods are illustrated by application to preliminary data from a study aimed at identifying hitherto unsuspected occupational carcinogens.

Bayes Theorem↗

The 'Ideal Homunculus': decoding neural population signals.

Information processing in the nervous system involves the activity of large populations of neurons. It is possible, however, to interpret the activity of relatively small numbers of cells in terms of meaningful aspects of the environment. 'Bayesian inference' provides a systematic and effective method of combining information from multiple cells to accomplish this. It is not a model of a neural mechanism (neither are alternative methods, such as the population vector approach) but a tool for analysing neural signals. It does not require difficult assumptions about the nature of the dimensions underlying cell selectivity, about the distribution and tuning of cell responses or about the way in which information is transmitted and processed. It can be applied to any parameter of neural activity (for example, firing rate or temporal pattern). In this review, we demonstrate the power of Bayesian analysis using examples of visual responses of neurons in primary visual and temporal cortices. We show that interaction between correlation in mean responses to different stimuli (signal) and correlation in response variability within stimuli (noise) can lead to marked improvement of stimulus discrimination using population responses.

Animals↗

The contributions of Jerome Cornfield to the theory of statistics.

This paper is a review of the contributions of Jerome Cornfield to the theory of statistics. It discusses several highlights of his theoretical work as well as describing his philosophy relating theory to application. The three areas discussed are: linear programming, urn sampling and its generalizations to the analysis of variance, and Bayesian inference. It is not widely known that Jerome Cornfield was perhaps the first to formulate and approximately solve the linear programming problem in 1941. His formulation was made for the famous "Diet Problem". An early publication introduced the method of indicator random variables in the context of urn sampling. This simple method allowed straightforward calculations of the low order moments for estimates arising from sampling finite populations and was later generalized to the two-way analysis of variance. The application of the urn sampling model to the analysis of variance served to illuminate how one chooses proper error terms for making tests in the analysis of variance table. Jerome Cornfield's philosophy on applications of statistics was dominated by a Bayesian outlook. His theoretical contributions in the past two decades were mainly concerned with the development of Bayesian ideas and methods. A brief survey is made of his main contributions to this area. A particularly noteworthy result was his demonstration that for the two-sample slippage problem of location, the likelihood function under a permutation setting is uninformative for the slippage parameter. However, the posterior distribution differs from the prior distribution despite the fact that the likelihood is uninformative.

Bayes Theorem↗

Bayesian analysis of prevalence with covariates using simulation-based techniques: applications to HIV screening.

Ignoring the limited precision of medical diagnostic tests can incur serious bias in prevalence estimation. Conversely, treating the values of sensitivity and specificity as constants, as in most studies, inevitably underestimates the variability of prevalence estimates. Bayesian inference provides a natural framework with which to integrate the variability in the estimates of sensitivity and specificity with estimation of prevalence. However, the resulting model becomes quite complicated and presents a computational challenge. Recently, Mendoza-Blanco et al. proposed a missing-data approach with simulation-based techniques to deal with the computational difficulties. Although their approach is quite effective in reducing the computational complexity into manageable tasks, their developed methodology is not general enough for modelling the effects of covariates in prevalence estimation. In this paper, we extend their work in this direction by combining their missing-data approach with a latent variable technique for modelling discrete data. The present work also generalizes the methods of Albert and Chib for Bayesian analysis of binary response data with errors in the response. We illustrate the methodology with several real data examples extracted from the literature.

AIDS Serodiagnosis↗

Trial-to-trial variability of cortical evoked responses: implications for the analysis of functional connectivity.

OBJECTIVES: The time series of single trial cortical evoked potentials typically have a random appearance, and their trial-to-trial variability is commonly explained by a model in which random ongoing background noise activity is linearly combined with a stereotyped evoked response. In this paper, we demonstrate that more realistic models, incorporating amplitude and latency variability of the evoked response itself, can explain statistical properties of cortical potentials that have often been attributed to stimulus-related changes in functional connectivity or other intrinsic neural parameters. METHODS: Implications of trial-to-trial evoked potential variability for variance, power spectrum, and interdependence measures like cross-correlation and spectral coherence, are first derived analytically. These implications are then illustrated using model simulations and verified experimentally by the analysis of intracortical local field potentials recorded from monkeys performing a visual pattern discrimination task. To further investigate the effects of trial-to-trial variability on the aforementioned statistical measures, a Bayesian inference technique is used to separate single-trial evoked responses from the ongoing background activity. RESULTS: We show that, when the average event-related potential (AERP) is subtracted from single-trial local field potential time series, a stimulus phase-locked component remains in the residual time series, in stark contrast to the assumption of the common model that no such phase-locked component should exist. Two main consequences of this observation are demonstrated for statistical measures that are computed on the residual time series. First, even though the AERP has been subtracted, the power spectral density, computed as a function of time with a short sliding window, can nonetheless show signs of modulation by the AERP waveform. Second, if the residual time series of two channels co-vary, then their cross-correlation and spectral coherence time functions can also be modulated according to the shape of the AERP waveform. Bayesian estimation of single-trial evoked responses provides further proof that these time-dependent statistical changes are due to remnants of the evoked phase-locked component in the residual time series. CONCLUSIONS: Because trial-to-trial variability of the evoked response is commonly ignored as a contributing factor in evoked potential studies, stimulus-related modulations of power spectral density, cross-correlation, and spectral coherence measures is often attributed to dynamic changes of the connectivity within and among neural populations. This work demonstrates that trial-to-trial variability of the evoked response must be considered as a possible explanation of such modulation.

Animals↗

Plastome evolution and phylogenomic relationships in Ajuga (Lamiaceae, Ajugoideae).

BACKGROUND: Ajuga is currently known to include approximately 69 species, with a combined distribution extending throughout Eurasia, Africa, and Australia. Its popularity and significance are largely based on an extensive history of medicinal and horticultural use. It is divided into two sections based on morphological characters, and this sectional classification is also reflected in pronounced geographic patterns. Although previous studies have largely focused on Ajuga sect. Ajuga in East Asia, A. sect. Chamaepithys, which ranges from the Mediterranean to Central Asia, remains insufficiently sampled, thereby limiting a comprehensive understanding of infrageneric sectional relationships within the genus. Here, we generated complete plastid genomes for 12 species representing both sections of the genus and used these data to characterize plastome structure and infer evolutionary relationships. RESULTS: In this study, 21 Ajuga plastomes were analyzed, including 12 newly sequenced plastomes and 9 previously published plastomes representing 19 species. Comparative analyses showed that all plastomes exhibited a highly conserved quadripartite structure, with genome sizes ranging from 149,963 to 150,740 bp and GC contents varying from 38.2% to 38.3%. Each plastome contained 133 genes, including 88 protein-coding genes, 37 transfer RNA genes, and 8 ribosomal RNA genes. The boundaries between the inverted repeat (IR) and single-copy (SC) regions were also highly conserved across species. In addition, 796 simple sequence repeats (SSRs), 874 long repeat sequences (LRSs), and 12 highly variable regions (ccsA-ndhD, ndhF-rpl32, petA-psbJ, rpl32-trnL-UAG, rps2-rpoC2, trnH-GUG-psbA, trnK-UUU-rps16, trnP-UGG-psaJ, trnT-UGU-trnL-UAA, ycf15-trnL-CAA, ndhF, and ycf1) were identified among the 21 plastomes. Phylogenetic analyses based on four datasets and conducted using Maximum Likelihood and Bayesian Inference recovered two major clades corresponding to the traditionally recognized sectional classification, with one distributed from the Mediterranean to Central Asia and the other in East Asia. CONCLUSION: This study represents the most comprehensive plastome-based sampling of Ajuga to date, including representative species from the Mediterranean, Central Asia, and East Asia. Our results have significantly enhanced our understanding of its infrageneric relationships. The plastome resources generated in this study provide a valuable foundation for future research on species delimitation, phylogeny, and the evolutionary history of Ajuga.

Phylogeny↗

Evolutionary HMMs: a Bayesian approach to multiple alignment.

MOTIVATION: We review proposed syntheses of probabilistic sequence alignment, profiling and phylogeny. We develop a multiple alignment algorithm for Bayesian inference in the links model proposed by Thorne et al. (1991, J. Mol. Evol., 33, 114-124). The algorithm, described in detail in Section 3, samples from and/or maximizes the posterior distribution over multiple alignments for any number of DNA or protein sequences, conditioned on a phylogenetic tree. The individual sampling and maximization steps of the algorithm require no more computational resources than pairwise alignment. METHODS: We present a software implementation (Handel) of our algorithm and report test results on (i) simulated data sets and (ii) the structurally informed protein alignments of BAliBASE (Thompson et al., 1999, Nucleic Acids Res., 27, 2682-2690). RESULTS: We find that the mean sum-of-pairs score (a measure of residue-pair correspondence) for the BAliBASE alignments is only 13% lower for Handelthan for CLUSTALW(Thompson et al., 1994, Nucleic Acids Res., 22, 4673-4680), despite the relative simplicity of the links model (CLUSTALW uses affine gap scores and increased penalties for indels in hydrophobic regions). With reference to these benchmarks, we discuss potential improvements to the links model and implications for Bayesian multiple alignment and phylogenetic profiling. AVAILABILITY: The source code to Handelis freely distributed on the Internet at http://www.biowiki.org/Handel under the terms of the GNU Public License (GPL, 2000, http://www.fsf.org./copyleft/gpl.html).

Algorithms↗