Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Change-point analysis of neuron spike train data.

In many medical experiments, data are collected across time, over a number of similar trials, or over a number of experimental units. As is the case of neuron spike train studies, these data may be in the form of counts of events per unit of time. These counts may be correlated within each trial. It is often of interest to know if the introduction of an intervention, such as the application of a stimulus, affects the distribution of the counts over the course of the experiment. In such investigations, each trial generates a sequence of data that may or may not contain a change in distribution at some point in time. Each sequence of integer counts can be viewed as arising from a Poisson process and are therefore independently distributed or as an integer-valued time series that allows for correlations between these counts. The main aim of this paper is to show how the ensemble of sample paths may be used to make inference about the distribution of the instantaneous times of change in a given population. This will be accomplished using a Bayesian hierarchical model for these change-points in time. A bonus of these models is they also allow for inference about the probability of a change in each unit and the magnitude of the effects, if any. The use of such change-point models on integer-valued time series is illustrated on neuron spike train data, although the methods can be applied to other situations where integer-valued processes arise.

Action Potentials↗

[Bayesian thinking on its way into medical statistics?].

BACKGROUND: Bayesian statistical analysis is a paradigm quite different from traditional statistical inference. We wanted to show the usefulness of this approach for some medical problems. MATERIALS AND METHODS: We started with Bayes equation as it is used for estimating the probability of illness based on a specific laboratory test. We also looked into a recent Cochrane report on mammography that accepted two studies as valid and five others as biased. In comparison we used examples of clinical trials from other areas that have been misinterpreted by the use of a traditional statistical approach only. RESULTS: We found that by taking into account our prior beliefs about the likely effects on breast cancer mortality of routine radiological screening programmes, the new data fit well into an estimate of a 5% mortality reduction with a 77% chance that there is a positive effect of screening. INTERPRETATION: Bayesian statistics is helpful in making decisions on the basis of experimental evidence by taking into account our prior knowledge, whereas p-values in traditional statistics only give information on how often we will end up with a false positive conclusion in the long run.

Bayes Theorem↗

Robust Bayesian methods for monitoring clinical trials.

Bayesian methods for the analysis of clinical trials data have received increasing attention recently as they offer an approach for dealing with difficult problems that arise in practice. A major criticism of the Bayesian approach, however, has focused on the need to specify a single, often subjective, prior distribution for the parameters of interest. In an attempt to address this criticism, we describe methods for assessing the robustness of the posterior distribution to the specification of the prior. The robust Bayesian approach to data analysis replaces the prior distribution with a class of prior distributions and investigates how the inferences might change as the prior varies over this class. The purpose of this paper is to illustrate the application of robust Bayesian methods to the analysis of clinical trials data. Using two examples of clinical trials taken from the literature, we illustrate how to use these methods to help a data monitoring committee decide whether or not to stop a trial early.

Bayes Theorem↗

Bayesian reconstruction and differential testing of excised introns.

MOTIVATION: Characterizing the differential excision of introns is critical for understanding the functional complexity of a cell or tissue, from normal developmental processes to disease pathogenesis. Most transcript reconstruction methods infer full-length transcripts from high-throughput sequencing data. However, this is a challenging task due to incomplete annotations and the heterogeneous expression of transcripts across cell-types, tissues, and experimental conditions. Several recent methods circumvent these difficulties by considering local splicing events, but these methods lose transcript-level splicing information and may conflate similar, but distinct transcripts. RESULTS: In this work, we formalize a new transcript reconstruction problem that interpolates between the full-length and local splicing perspectives by considering sequences of exon-exon junctions (SEEJs) that co-occur in transcripts. We then present a hierarchical Bayesian admixture model and posterior inference algorithms for computing SEEJs (BSEEJ), and a generalized linear model for characterizing differential SEEJ usage based on model parameter estimates. We show that BSEEJ achieves high F1 score for reconstruction tasks and improved accuracy and sensitivity in differential splicing when compared with six transcript and local splicing methods on simulated data. Lastly, we evaluate BSEEJ on experimental data based on transcript reconstruction, novelty of transcripts produced, model sensitivity to hyperparameters, and a functional analysis of differentially expressed SEEJs. AVAILABILITY AND IMPLEMENTATION: BSEEJ is freely available at https://github.com/bayesomicslab/BSEEJ.

Bayes Theorem↗

Modeling cellular processes with variational Bayesian cooperative vector quantizer.

Gene expression of a cell is controlled by sophisticated cellular processes. The capability of inferring the states of these cellular processes would provide insight into the mechanism of gene expression control system. In this paper, we propose and investigate the cooperative vector quantizer (CVQ) model for analysis of microarray data. The CVQ model could be capable of decomposing observed microarray data into many different regulatory subprocesses. To make the CVQ analysis tractable we develop and apply variational approximations. Bayesian model selection is employed in the model, so that the optimal number processes is determined purely from observed micro-array data. We test the model and algorithms on two datasets: (1) simulated gene-expression data and (2) real-world yeast cell-cycle microarray data. The results illustrate the ability of the CVQ approach to recover and characterize regulatory gene expression subprocesses, indicating a potential for advanced gene expression data analysis.

Algorithms↗

Alternatives to statistical hypothesis testing in ecology: a guide to self teaching.

Statistical methods emphasizing formal hypothesis testing have dominated the analyses used by ecologists to gain insight from data. Here, we review alternatives to hypothesis testing including techniques for parameter estimation and model selection using likelihood and Bayesian techniques. These methods emphasize evaluation of weight of evidence for multiple hypotheses, multimodel inference, and use of prior information in analysis. We provide a tutorial for maximum likelihood estimation of model parameters and model selection using information theoretics, including a brief treatment of procedures for model comparison, model averaging, and use of data from multiple sources. We discuss the advantages of likelihood estimation, Bayesian analysis, and meta-analysis as ways to accumulate understanding across multiple studies. These statistical methods hold promise for new insight in ecology by encouraging thoughtful model building as part of inquiry, providing a unified framework for the empirical analysis of theoretical models, and by facilitating the formal accumulation of evidence bearing on fundamental questions.

Algorithms↗

Bayesian eggs and Bayesian omelettes: reply to Stern (2005).

In this response to Stern's (2005) discussion of Klugkist, Laudy, and Hoijtink (2005), model inference based on posterior probabilities on the parameter space is discussed. Furthermore, the authors respond to Stern's example in which all possible orderings are included via a short discussion of exploratory versus theory-based modeling. Finally, the authors show that the Bayesian approach is flexible and can deal with many types of constraints. This is illustrated using a model with constraints on the differences between means.

Bayes Theorem↗

Shared genetic architecture of obesity and gastroesophageal reflux disease.

Obesity is identified as a risk factor of gastroesophageal reflux disease (GERD). This study aims to elucidate the shared genetic architecture of obesity-related phenotypes and GERD. Based on the publicly available genome-wide association studies' datasets, this genome-wide pleiotropic association study was conducted with various genetic approaches (including linkage disequilibrium score regression, high-definition likelihood inference for genetic correlations, pleiotropic analysis under composite null hypothesis, Functional Mapping and Annotation, Bayesian colocalization, summary-based Mendelian randomization, and multi-marker analysis of genomic annotation analysis) sequentially to unravel the genetic associations from single-nucleotide polymorphism to gene levels, and to reveal the underlying shared genetic architecture between obesity-related phenotypes and GERD. This study discovered shared genetic mechanisms between GERD and several obesity-related phenotypes, including arm fat percentage (left), arm fat percentage (right), leg fat percentage (left), leg fat percentage (right), trunk fat percentage, waist-to-hip ratio, and body mass index. Significant genetic correlations were observed by linkage disequilibrium score regression and high-definition likelihood inference for genetic correlations, with multiple associated pleiotropic loci and their mapped genes identified by pleiotropic analysis under composite null hypothesis, Functional Mapping and Annotation, Bayesian colocalization, summary-based Mendelian randomization, and multi-marker analysis of genomic annotation analysis. Additionally, several brain tissues were identified to be linked to both obesity and GERD by multi-marker analysis of genomic annotation. This research provided strong evidence of genetic correlations and brought novel insights into the underlying genetic connections and shared genetic architectures of obesity and GERD.

Humans↗

Efficient temporal probabilistic reasoning via context-sensitive model construction.

We present a language for representing context-sensitive temporal probabilistic knowledge. Context constraints allow inference to be focused on only the relevant portions of the probabilistic knowledge. We provide a declarative semantics for our language. We present a sound and complete algorithm for computing posterior probabilities of temporal queries, as well as an efficient implementation of the algorithm. Throughout we illustrate the approach with the problem of reasoning about the effects of medications and interventions on the state of a patient in cardiac arrest. We empirically evaluate the efficiency of our system by comparing its inference times on problems in this domain with those of standard Bayesian network representations of the problems.

Algorithms↗

Withdrawal time estimation of veterinary drugs: extending the range of statistical methods.

In order to use a drug in a food producing animal, evidence has to be provided that after a certain withdrawal time, drug residues in tissues, such as muscle meat, fat, liver, kidney etc., are below a given maximum residue limit (MRL), for a majority of animals. Several statistical methods, both regression based and nonparametric based methods, have been proposed, each relying on different sets of assumptions, which may or may not hold for the specific data situation. The purpose of this paper is to enrich the range of methods, i.e. to provide approaches for situations where current methods are inappropriate. Bayesian methods, using Markov chain Monte Carlo, are proposed to derive inference on the parameters of interest.

Animals↗

On marker-assisted prediction of genetic value: beyond the ridge.

Marked-assisted genetic improvement of agricultural species exploits statistical dependencies in the joint distribution of marker genotypes and quantitative traits. An issue is how molecular (e.g., dense marker maps) and phenotypic information (e.g., some measure of yield in plants) is to be used for predicting the genetic value of candidates for selection. Multiple regression, selection index techniques, best linear unbiased prediction, and ridge regression of phenotypes on marker genotypes have been suggested, as well as more elaborate methods. Here, phenotype-marker associations are modeled hierarchically via multilevel models including chromosomal effects, a spatial covariance of marked effects within chromosomes, background genetic variability, and family heterogeneity. Lorenz curves and Gini coefficients are suggested for assessing the inequality of the contribution of different marked effects to genetic variability. Classical and Bayesian methods are presented. The Bayesian approach includes a Markov chain Monte Carlo implementation. The generality and flexibility of the Bayesian method is illustrated when a Lorenz curve is to be inferred.

Animals↗

Evolutionary dynamics of human retroviruses investigated through full-genome scanning.

To test hypotheses on the differences in retroviral genetic diversity, we compared the evolutionary dynamics of the human immunodeficiency virus type 1 (HIV-1) group M and the primate T-cell lymphotropic virus (PTLV) using a full-genome analysis. Evolutionary rates and nonsynonymous/synonymous substitution rate ratios were estimated across the genome using a maximum likelihood sliding window approach, and molecular clock properties were investigated. We confirm a remarkable difference in genetic stability and selective pressure at the interhost level. While there is evidence for adaptive evolution in HIV-1, the evolution of PTLV is almost exclusively characterized by negative selection or nearly neutral processes. For both retroviruses, evolutionary rate estimates across the genome reflect the differential selective constraints. However, based on the relationship between evolutionary rate and selective pressure and based on the comparison of synonymous substitution rates, the differences in rate between HIV-1 and PTLV cannot be explained by selective forces only. Several evolutionary and statistical assumptions, examined using a Bayesian coalescent method, were shown to have little influence on our inference.

Bayes Theorem↗

An extended general location model for causal inferences from data subject to noncompliance and missing values.

Noncompliance is a common problem in experiments involving randomized assignment of treatments, and standard analyses based on intention-to-treat or treatment received have limitations. An attractive alternative is to estimate the Complier-Average Causal Effect (CACE), which is the average treatment effect for the subpopulation of subjects who would comply under either treatment (Angrist, Imbens, and Rubin, 1996, Journal of American Statistical Association 91, 444-472). We propose an extended general location model to estimate the CACE from data with noncompliance and missing data in the outcome and in baseline covariates. Models for both continuous and categorical outcomes and ignorable and latent ignorable (Frangakis and Rubin, 1999, Biometrika 86, 365-379) missing-data mechanisms are developed. Inferences for the models are based on the EM algorithm and Bayesian MCMC methods. We present results from simulations that investigate sensitivity to model assumptions and the influence of missing-data mechanism. We also apply the method to the data from a job search intervention for unemployed workers.

Algorithms↗

Habitat barriers limit gene flow and illuminate historical events in a wide-ranging carnivore, the American puma.

We examined the effects of habitat discontinuities on gene flow among puma (Puma concolor) populations across the southwestern USA. Using 16 microsatellite loci, we genotyped 540 pumas sampled throughout the states of Utah, Colorado, Arizona, and New Mexico, where a high degree of habitat heterogeneity provides for a wide range of connective habitat configurations between subpopulations. We investigated genetic structuring using complementary individual- and population-based analyses, the latter employing a novel technique to geographically cluster individuals without introducing investigator bias. The analyses revealed genetic structuring at two distinct scales. First, strikingly strong differentiation between northern and southern regions within the study area suggests little migration between them. Second, within each region, gene flow appears to be strongly limited by distance, particularly in the presence of habitat barriers such as open desert and grasslands. Northern pumas showed both reduced genetic diversity and greater divergence from a hypothetical ancestral population based on Bayesian clustering analyses, possibly reflecting a post-Pleistocene range expansion. Bayesian clustering results were sensitive to sampling density, which may complicate inference of numbers of populations when using this method. The results presented here build on those of previous studies, and begin to complete a picture of how different habitat types facilitate or impede gene flow among puma populations.

Animals↗

Introgression in natural populations of bioindicators: a case study of Carabus splendens and Carabus punctatoauratus.

The evolutionary importance of hybridization in wild plants and animals has become increasingly widely recognized in the last decade. In practical terms, hybridization provides an exceptionally tough set of problems for conservation biologists. We illustrate this in a case study of two Carabidae species widely used to evaluate the impact of human activities on biodiversity. These two species live in a complex mosaic of sympatry/allopatry and are known to hybridize in controlled conditions. Hybridization has not been quantified in natural populations to date due to the lack of a simple set of phenotypic traits for identifying hybrids. We thus screened for hybrids in natural populations, by multilocus genotyping at nine microsatellite loci. A high level of genetic differentiation between these two taxa was observed, as shown by allelic frequency distributions. Two Bayesian assignment procedures without obligatory pure taxon references were used to infer different classes of hybrids (F(1), F(2) and backcrosses) and mixture proportions between the two species. A low level of hybridization (F(1) genotypes) was observed in natural populations, contrasting with results obtained in controlled conditions. A high level of introgression was, however, detected at three of 12 sites, as revealed by the detection of backcrossed genotypes. This interspecific gene flow was detected in a limited zone of the common geographical range of the two species and was not related to the pattern of sympatry/allopatry. We then considered the origin and repercussions of this introgression, based on intraspecific genetic diversity and geographical structure.

Animals↗

Statistical methods for geographical surveillance in veterinary epidemiology.

Spatial clustering and cluster detection are statistical analysis developed to address relevant scientific hypothesis. The difficulty stays in the large number of alternative hypothesis due to the different mechanisms that could generate the anomalous cases aggregation. We review methods for marked point data (case/control) aimed to describe spatial intensity of disease risk, to test for randomness and to locate significant excesses. Bayesian Gaussian Spatial Exponential models are used to illustrate probabilistic aspects and the link with simpler non parametric tools are shown. We develop an informal guideline to the analysis and used data on faecal contamination and dog parasitic diseases in the city of Naples, Italy. Kernel density estimation resulted very sensitive to bandwidth choice and overemphasized localized excess, Ripley'K function and Cuzick-Edwards test were very consistent each other while the SatScan failed to detect excesses. The spatial range was around 600 meters and justifies several small clusters. Bayesian models were very powerful in reconstructing the phenomenon and allow inference on model parameters in good agreement with the non parametric analysis.

Algorithms↗

An empirical Bayesian significance test of cDNA library data.

Automated high-throughput sequencing of cDNA clones from numerous libraries has generated a wealth of information about both genome sequence and relative transcript abundances. A common statistical challenge in the analysis of library sequences is to infer whether there is differential expression for the same transcript under two different conditions, such as normal and diseased tissue. In contrast to the continuously variable intensity measurements from microarray experiments, data from cDNA library sequencing presents itself as a discrete count of the incidence of some clone or transcript in a finite sample. In this paper, we first propose a statistical model for data generated from cDNA library sequencing efforts. The model is based on the Poisson mixed with generalized inverse Gaussian (PGIG), introduced by Sichel (1971, 1975). PGIG has been used in modeling population abundance, ecological studies, word frequencies in publications, etc. Using data from the literature, we show that the proposed model provides a good fit to the observed data. Using this new model for cDNA library data, we developed an empirical Bayesian significance test (EBST) for inferring the statistical significance of differential gene expression from discrete data.

Bayes Theorem↗

Model-based assignment and inference of protein backbone Nuclear Magnetic Resonances.

Nuclear Magnetic Resonance (NMR) spectroscopy is a key experimental technique used to study protein structure, dynamics, and interactions. NMR methods face the bottleneck of spectral analysis, in particular determining the resonance assignments, which help define the mapping between atoms in the protein and peaks in the spectra. A substantial amount of noise in spectral data, along with ambiguities in interpretation, make this analysis a daunting task, and there exists no generally accepted measure of uncertainty associated with the resulting solutions. This paper develops a model-based inference approach that addresses the problem of characterizing uncertainty in backbone resonance assignment. We argue that NMR spectra are subject to random variation, and ignoring this stochasticity can lead to false optimism and erroneous conclusions. We propose a Bayesian statistical model that accounts for various sources of uncertainty and provides an automatable framework for inference. While assignment has previously been viewed as a deterministic optimization problem, we demonstrate the importance of considering all solutions consistent with the data, and develop an algorithm to search this space within our statistical framework. Our approach is able to characterize the uncertainty associated with backbone resonance assignment in several ways: 1) it quantifies of uncertainty in the individually assigned resonances in terms of their posterior standard deviations; 2) it assesses the information content in the data with a posterior distribution of plausible assignments; and 3) it provides a measure of the overall plausibility of assignments. We demonstrate the value of our approach in a study of experimental data from two proteins, Human Ubiquitin and Cold-shock protein A from E. coli. In addition, we provide simulations showing the impact of experimental conditions on uncertainty in the assignments.

Journal Article↗