Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Incorporating biological knowledge into evaluation of causal regulatory hypotheses.

Biological data can be scarce and costly to obtain. The small number of samples available typically limits statistical power and makes reliable inference of causal relations extremely difficult. However, we argue that statistical power can be increased substantially by incorporating prior knowledge and data from diverse sources. We present a Bayesian framework that combines information from different sources and we show empirically that this lets one make correct causal inferences with small sample sizes that otherwise would be impossible.

Algorithms↗

A bayesian analysis of metazoan mitochondrial genome arrangements.

Genome arrangements are a potentially powerful source of information to infer evolutionary relationships among distantly related taxa. Mitochondrial genome arrangements may be especially informative about metazoan evolutionary relationships because (1) nearly all animals have the same set of definitively homologous mitochondrial genes, (2) mitochondrial genome rearrangement events are rare relative to changes in sequences, and (3) the number of possible mitochondrial genome arrangements is huge, making convergent evolution of genome arrangements appear highly unlikely. In previous studies, phylogenetic evidence in genome arrangement data is nearly always used in a qualitative fashion-the support in favor of clades with similar or identical genome arrangements is considered to be quite strong, but is not quantified. The purpose of this article is to quantify the uncertainty among the relationships of metazoan phyla on the basis of mitochondrial genome arrangements while incorporating prior knowledge of the monophyly of various groups from other sources. The work we present here differs from our previous work in the statistics literature in that (1) we incorporate prior information on classifications of metazoans at the phylum level, (2) we describe several advances in our computational approach, and (3) we analyze a much larger data set (87 taxa) that consists of each unique, complete mitochondrial genome arrangement with a full complement of 37 genes that were present in the NCBI (National Center for Biotechnology Information) database at a recent date. In addition, we analyze a subset of 28 of these 87 taxa for which the non-tRNA mitochondrial genomes are unique where the assumption of our inversion-only model of rearrangement is more plausible. We present summaries of Bayesian posterior distributions of tree topology on the basis of these two data sets.

Animals↗

A heuristic Bayesian method for segmenting DNA sequence alignments and detecting evidence for recombination and gene conversion.

We propose a heuristic approach to the detection of evidence for recombination and gene conversion in multiple DNA sequence alignments. The proposed method consists of two stages. In the first stage, a sliding window is moved along the DNA sequence alignment, and phylogenetic trees are sampled from the conditional posterior distribution with MCMC. To reduce the noise intrinsic to inference from the limited amount of data available in the typically short sliding window, a clustering algorithm based on the Robinson-Foulds distance is applied to the trees thus sampled, and the posterior distribution over tree clusters is obtained for each window position. While changes in this posterior distribution are indicative of recombination or gene conversion events, it is difficult to decide when such a change is statistically significant. This problem is addressed in the second stage of the proposed algorithm, where the distributions obtained in the first stage are post-processed with a Bayesian hidden Markov model (HMM). The emission states of the HMM are associated with posterior distributions over phylogenetic tree topology clusters. The hidden states of the HMM indicate putative recombinant segments. Inference is done in a Bayesian sense, sampling parameters from the posterior distribution with MCMC. Of particular interest is the determination of the number of hidden states as an indication of the number of putative recombinant regions. To this end, we apply reversible jump MCMC, and sample the number of hidden states from the respective posterior distribution.

Actins↗

Bayesian analyses of admixture in wild and domestic cats (Felis silvestris) using linked microsatellite loci.

Methods recently developed to infer population structure and admixture mostly use individual genotypes described by unlinked neutral markers. However, Hardy-Weinberg and linkage disequilibria among independent markers decline rapidly with admixture time, and the admixture signals could be lost in a few generations. In this study, we aimed to describe genetic admixture in 182 European wild and domestic cats (Felis silvestris), which hybridize sporadically in Italy and extensively in Hungary. Cats were genotyped at 27 microsatellites, including 21 linked loci mapping on five distinct feline linkage groups. Genotypes were analysed with structure 2.1, a Bayesian procedure designed to model admixture linkage disequilibrium, which promises to assess efficiently older admixture events using tightly linked markers. Results showed that domestic and wild cats sampled in Italy were split into two distinct clusters with average proportions of membership Q > 0.90, congruent with prior morphological identifications. In contrast, free-living cats sampled in Hungary were assigned partly to the domestic and the wild cat clusters, with Q < 0.50. Admixture analyses of individual genotypes identified, respectively, 5/61 (8%), and 16-20/65 (25-31%) hybrids among the Italian wildcats and Hungarian free-living cats. Similar results were obtained in the past using unlinked loci, although the new linked markers identified additional admixed wildcats in Italy. Linkage analyses confirm that hybridization is limited in Italian, but widespread in Hungarian wildcats, a population that is threatened by cross-breeding with free-ranging domestic cats. The total panel of 27 loci performed better than the linked loci alone in the identification of domestic and known hybrid cats, suggesting that a large number of linked plus unlinked markers can improve the results of admixture analyses. Inferred recombination events led to identify the population of origin of chromosomal segments, suggesting that admixture mapping experiments can be designed also in wild populations.

Animals↗

A Bayesian approach to reconstructing genetic regulatory networks with hidden factors.

MOTIVATION: We have used state-space models (SSMs) to reverse engineer transcriptional networks from highly replicated gene expression profiling time series data obtained from a well-established model of T cell activation. SSMs are a class of dynamic Bayesian networks in which the observed measurements depend on some hidden state variables that evolve according to Markovian dynamics. These hidden variables can capture effects that cannot be directly measured in a gene expression profiling experiment, for example: genes that have not been included in the microarray, levels of regulatory proteins, the effects of mRNA and protein degradation, etc. RESULTS: We have approached the problem of inferring the model structure of these state-space models using both classical and Bayesian methods. In our previous work, a bootstrap procedure was used to derive classical confidence intervals for parameters representing 'gene-gene' interactions over time. In this article, variational approximations are used to perform the analogous model selection task in the Bayesian context. Certain interactions are present in both the classical and the Bayesian analyses of these regulatory networks. The resulting models place JunB and JunD at the centre of the mechanisms that control apoptosis and proliferation. These mechanisms are key for clonal expansion and for controlling the long term behavior (e.g. programmed cell death) of these cells. AVAILABILITY: Supplementary data is available at http://public.kgi.edu/wild/index.htm and Matlab source code for variational Bayesian learning of SSMs is available at http://www.cse.ebuffalo.edu/faculty/mbeal/software.html.

Bayes Theorem↗

On some applications of Bayesian methods in cancer clinical trials.

The NCCTG randomized controlled clinical trial for the treatment of advanced colorectal carcinoma is a wonderful case study of the dynamic interplay between scientific learning and statistical inference. Ethical concerns for minimizing the number of patients assigned to an inferior treatment and interest in identifying subsets of patients for whom a treatment is most likely efficacious pose challenging problems for the practice of statistics. In the first part of this paper, I comment on the applications of Bayesian methods to these problems in the NCCTG trial as presented by Freedman and Spieglehalter and Dixon and Simon, respectively. In the second part of this paper, I discuss and illustrate a Bayesian approach to model sensitivity analysis with a particular focus on model specification and criticism. The Bayesian approach provides a formal methodology to assess the sensitivity of inferences to the inputs into an analysis so that it is possible to investigate the consequences of the specification of the model. I apply these methods to the specification and criticism of a class of survival models for the analysis of survival times in the NCCTG trial.

Antineoplastic Combined Chemotherapy Protocols↗

Inferring gene regulatory networks with time delays using a genetic algorithm.

Recently a state-space model with time delays for inferring gene regulatory networks was proposed. It was assumed that each regulation between two internal state variables had multiple time delays. This assumption caused underestimation of the model with many current gene expression datasets. In biological reality, one regulatory relationship may have just a single time delay, and not multiple time delays. This study employs Boolean variables to capture the existence of the time-delayed regulatory relationships in gene regulatory networks in terms of the state-space model. As the solution space of time delayed relationships is too large for an exhaustive search, a genetic algorithm (GA) is proposed to determine the optimal Boolean variables (the optimal time-delayed regulatory relationships). Coupled with the proposed GA, Bayesian information criterion (BIC) and probabilistic principle component analysis (PPCA) are employed to infer gene regulatory networks with time delays. Computational experiments are performed on two real gene expression datasets. The results show that the GA is effective at finding time-delayed regulatory relationships. Moreover, the inferred gene regulatory networks with time delays from the datasets improve the prediction accuracy and possess more of the expected properties of a real network, compared to a gene regulatory network without time delays.

Algorithms↗

The irrelevance of inference: a decision-making approach to the stochastic evaluation of health care technologies.

The literature which considers the statistical properties of cost-effectiveness analysis has focused on estimating the sampling distribution of either an incremental cost-effectiveness ratio or incremental net benefit for classical inference. However, it is argued here that rules of inference are arbitrary and entirely irrelevant to the decisions which clinical and economic evaluations claim to inform. Decisions should be based only on the mean net benefits irrespective of whether differences are statistically significant or fall outside a Bayesian range of equivalence. Failure to make decisions in this way by accepting the arbitrary rules of inference will impose costs which can be measured in terms of resources or health benefits forgone. The distribution of net benefit is only relevant to deciding whether more information is required. A framework for decision making and establishing the value of additional information is presented which is consistent with the decision rules in CEA. This framework can distinguish the simultaneous but conceptually separate steps of deciding which alternatives should be chosen, given existing information, from the question of whether more information should be acquired. It also ensures that the type of information acquired is driven by the objectives of the health care system, is consistent with the budget constraint on service provision and that research is designed efficiently.

Bayes Theorem↗

A Bayesian framework for multivariate differential analysis.

Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.

Bayes Theorem↗

Adaptive BCI based on variational Bayesian Kalman filtering: an empirical evaluation.

This paper proposes the use of variational Kalman filtering as an inference technique for adaptive classification in a brain computer interface (BCI). The proposed algorithm translates electroencephalogram segments adaptively into probabilities of cognitive states. It, thus, allows for nonstationarities in the joint process over cognitive state and generated EEG which may occur during a consecutive number of trials. Nonstationarities may have technical reasons (e.g., changes in impedance between scalp and electrodes) or be caused by learning effects in subjects. We compare the performance of the proposed method against an equivalent static classifier by estimating the generalization accuracy and the bit rate of the BCI. Using data from two studies with healthy subjects, we conclude that adaptive classification significantly improves BCI performance. Averaging over all subjects that participated in the respective study, we obtain, depending on the cognitive task pairing, an increase both in generalization accuracy and bit rate of up to 8%. We may, thus, conclude that adaptive inference can play a significant contribution in the quest of increasing bit rates and robustness of current BCI technology. This is especially true since the proposed algorithm can be applied in real time.

Algorithms↗

Bayesian integrated functional analysis of microarray data.

MOTIVATION: The statistical analysis of microarray data usually proceeds in a sequential manner, with the output of the previous step always serving as the input of the next one. However, the methods currently used in such analyses do not properly account for the fact that the intermediate results may not always be correct, then leading to cumulating error in the inferences drawn based on such steps. RESULTS: Here we show that, by an application of hierarchical Bayesian methodology, this sequential procedure can be replaced by a single joint analysis, while systematically accounting for the uncertainties in this process. Moreover, we can also integrate relevant functional information available from databases into such an analysis, thereby increasing the reliability of the biological conclusions that are drawn. We illustrate these points by analysing real data and by showing that the genes can be divided into categories of interest, with the defining characteristic depending on the biological question that is considered. We contend that the proposed method has advantages at two levels. First, there are gains in the statistical and biological results from the analysis of this particular dataset. Second, it opens up new possibilities in analysing microarray data in general.

Algorithms↗

Conflicting phylogenies of balsaminoid families and the polytomy in Ericales: combining data in a Bayesian framework.

The balsaminoid Ericales, namely Balsaminaceae, Marcgraviaceae, Tetrameristaceae, and Pellicieraceae have been confidently placed at the base of Ericales, but the relations among these families have been resolved differently in recent analyses. Sister to this basal group is a large polytomy comprising all other families of Ericales, which is associated with short internodes. Because there are more than 13 kb of sequences for a large sampling of representatives, a thorough examination of the available data with novel methods seemed in place. Because of its computational speed, Bayesian phylogenetics allows for the use of parameter-rich models that can accommodate differences in the evolutionary process between partitions in a simultaneous analysis. In addition, there are recently proposed Bayesian strategies of assessing incongruence between partitions. We have applied these methods to the current problems in Ericales phylogeny, taking into account reported pitfalls in Bayesian analysis such as model selection uncertainty. Based on our results we infer several, previously unresolved relationships in the order Ericales. In balsaminoid families, we find that the closest relatives of Balsaminaceae are Marcgraviaceae. In the Ericales polytomy, we find strong support for Pentaphylacaceae sensu APG II as the sister group of Maesaceae. In addition, Symplocaceae receive a position as sister to Theaceae and these families form a monophyletic group together with Styracaceae-Diapensiaceae. At the base of this clade are Actinidiaceae and Clethraceae. The positions of Ebenaceae and Lecythidaceae remain uncertain.

Balsaminaceae↗

Null hypothesis significance testing. On the survival of a flawed method.

Null hypothesis significance testing (NHST) is the researcher's workhorse for making inductive inferences. This method has often been challenged, has occasionally been defended, and has persistently been used through most of the history of scientific psychology. This article reviews both the criticisms of NHST and the arguments brought to its defense. The review shows that the criticisms address the logical validity of inferences arising from NHST, whereas the defenses stress the pragmatic value of these inferences. The author suggests that both critics and apologists implicitly rely on Bayesian assumptions. When these assumptions are made explicit, the primary challenge for NHST--and any system of induction--can be confronted. The challenge is to find a solution to the question of replicability.

Humans↗

Estimating absolute rates of synonymous and nonsynonymous nucleotide substitution in order to characterize natural selection and date species divergences.

The rate of molecular evolution can vary among lineages. Sources of this variation have differential effects on synonymous and nonsynonymous substitution rates. Changes in effective population size or patterns of natural selection will mainly alter nonsynonymous substitution rates. Changes in generation length or mutation rates are likely to have an impact on both synonymous and nonsynonymous substitution rates. By comparing changes in synonymous and nonsynonymous rates, the relative contributions of the driving forces of evolution can be better characterized. Here, we introduce a procedure for estimating the chronological rates of synonymous and nonsynonymous substitutions on the branches of an evolutionary tree. Because the widely used ratio of nonsynonymous and synonymous rates is not designed to detect simultaneous increases or simultaneous decreases in synonymous and nonsynonymous rates, the estimation of these rates rather than their ratio can improve characterization of the evolutionary process. With our Bayesian approach, we analyze cytochrome oxidase subunit I evolution in primates and infer that nonsynonymous rates have a greater tendency to change over time than do synonymous rates. Our analysis of these data also suggests that rates have been positively correlated.

Animals↗

Bayesian point estimation of quantitative trait loci.

In this article, we consider the problem of the estimation of quantitative trait loci (QTL), those chromosomal regions at which genetic information affecting some quantitative trait is encoded. Generally the number of such encoding sites is unknown, and associations between neutral molecular marker genotypes and observed trait phenotypes are sought to locate them. We consider a Bayesian model for simple experimental designs, and discuss the existing approaches to inference for this problem. In particular, we focus on locating positions of the best candidate markers segregating for the trait, a situation which is of primary interest in comparative mapping. We introduce a loss function for estimating both the number of QTL and their location, and we illustrate its application via simulated and real data.

Animals↗

Quantitative genetic models for describing simultaneous and recursive relationships between phenotypes.

Multivariate models are of great importance in theoretical and applied quantitative genetics. We extend quantitative genetic theory to accommodate situations in which there is linear feedback or recursiveness between the phenotypes involved in a multivariate system, assuming an infinitesimal, additive, model of inheritance. It is shown that structural parameters defining a simultaneous or recursive system have a bearing on the interpretation of quantitative genetic parameter estimates (e.g., heritability, offspring-parent regression, genetic correlation) when such features are ignored. Matrix representations are given for treating a plethora of feedback-recursive situations. The likelihood function is derived, assuming multivariate normality, and results from econometric theory for parameter identification are adapted to a quantitative genetic setting. A Bayesian treatment with a Markov chain Monte Carlo implementation is suggested for inference and developed. When the system is fully recursive, all conditional posterior distributions are in closed form, so Gibbs sampling is straightforward. If there is feedback, a Metropolis step may be embedded for sampling the structural parameters, since their conditional distributions are unknown. Extensions of the model to discrete random variables and to nonlinear relationships between phenotypes are discussed.

Bayes Theorem↗

A spatial statistical model for landscape genetics.

Landscape genetics is a new discipline that aims to provide information on how landscape and environmental features influence population genetic structure. The first key step of landscape genetics is the spatial detection and location of genetic discontinuities between populations. However, efficient methods for achieving this task are lacking. In this article, we first clarify what is conceptually involved in the spatial modeling of genetic data. Then we describe a Bayesian model implemented in a Markov chain Monte Carlo scheme that allows inference of the location of such genetic discontinuities from individual geo-referenced multilocus genotypes, without a priori knowledge on populational units and limits. In this method, the global set of sampled individuals is modeled as a spatial mixture of panmictic populations, and the spatial organization of populations is modeled through the colored Voronoi tessellation. In addition to spatially locating genetic discontinuities, the method quantifies the amount of spatial dependence in the data set, estimates the number of populations in the studied area, assigns individuals to their population of origin, and detects individual migrants between populations, while taking into account uncertainty on the location of sampled individuals. The performance of the method is evaluated through the analysis of simulated data sets. Results show good performances for standard data sets (e.g., 100 individuals genotyped at 10 loci with 10 alleles per locus), with high but also low levels of population differentiation (e.g., FST<0.05). The method is then applied to a set of 88 individuals of wolverines (Gulo gulo) sampled in the northwestern United States and genotyped at 10 microsatellites.

Animals↗

Bayesian selection of predictors of conception probabilities across the menstrual cycle.

There is increasing interest in identifying predictors of human fertility, including environmental exposures, behavioural factors, and biomarkers, such as mucus or reproductive hormones. Epidemiological studies typically measure fecundability, the per menstrual cycle probability of conception, using time to pregnancy data. A critical predictor, which is often ignored in the design or analysis, is the timing of non-contracepting intercourse in the menstrual cycle. In order to limit confounding by behavioural differences between exposure groups, it may be preferable to base inferences on day-specific conception probabilities in relation to intercourse timing. This article proposes Bayesian methods for selection of predictors of day-specific conception probabilities. A particular focus is the case in which data on ovulation timing are not available. We focus on the selection of fertile days in the cycle during which conception probabilities are non-negligible and predictors may play a role. Data from recent European and Italian prospective studies of daily fecundability are presented, and the proposed approach is used to estimate cervical mucus effects within a mid-cycle potentially fertile window using data from the Italian study.

Bayes Theorem↗