Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

A Markov chain model for animal estrous cycling data.

Estrous cycling data contain sequences of characters (e.g., DPEMD). Each sequence represents an animal's estrous cycle, with each character indicating the daily estrous cycle stage. Changes in the estrous cycle pattern, which is determined by estrous stage lengths, can provide information on adverse events. Stage lengths are not directly observable. However interval censored lengths for all but the first and the last stages in a sequence can be extracted from the data. We propose a Markov chain model to approximate the estrous cycling process. The transition probabilities from one stage to another can be derived by conditioning on stage lengths. Assuming Weibull distribution for stage lengths, with the second Weibull parameter depending upon treatment effects and animal-specific random effects, regression models on censored stage lengths are fitted. A Bayesian approach is used for inference on dose effects. The analysis is implemented with MCMC method in WinBUGS. An estrous cycling data set from a National Toxicology Program study is analyzed as an example.

Animals↗

Some approaches to the analysis of recurrent event data.

Methodological research in biostatistics has been dominated over the last twenty years by further development of Cox's regression model for life tables and of Nelder and Wedderburn's formulation of generalized linear models. In both of these areas the need to address the problems introduced by subject level heterogeneity has provided a major motivation, and the analysis of data concerning recurrent events has been widely discussed within both frameworks. This paper reviews this work, drawing together the parallel development of 'marginal' and 'conditional' approaches in survival analysis and in generalized linear models. Frailty models are shown to be a special case of a random effects generalization of generalized linear models, whereas marginal models for multivariate failure time data are more closely related to the generalized estimating equation approach to longitudinal generalized linear models. Computational methods for inference are discussed, including the Bayesian Markov chain Monte Carlo approach.

Algorithms↗

Elliptical selection experiment for the estimation of genetic parameters of the growth rate and feed conversion ratio in rabbits.

Two elliptical selection experiments were performed in two contemporary sire lines of rabbits (C and R) in order to optimize the experimental design for estimating the genetic parameters of the growth rate (GR) and feed conversion ratio (FCR). Twelve males and 19 females from line C, and 13 males and 23 females from line R, were selected from an ellipse defined by a quadratic index based on these traits. Data from 160 rabbits of each of the parental generations of lines C and R and their offspring (275 and 266 animals, respectively) were used for the analysis. A Bayesian framework was adopted for inference. Marginal posterior distributions of the genetic parameters were obtained by Gibbs sampling. An animal model including batch, parity order, litter size, and common environmental litter effects was assumed. Posterior means (posterior standard deviations) for heritabilities of GR and FCR were estimated to be 0.31 (0.10) and 0.31 (0.10), respectively, in line C and 0.21 (0.08) and 0.25 (0.12) in line R. Posterior means of the proportion of the variance due to common litter environmental effects were 0.14 (0.06) and 0.21 (0.06) for GR and FCR, respectively, in line C and 0.17 (0.06) and 0.22 (0.06) in line R. Posterior means of genetic correlation between both traits were -0.49 (0.25) in line C and -0.47 (0.32) in line R, indicating that selection for GR was expected to result in a similar correlated response in FCR in both lines.

Animal Feed↗

A hierarchical model for estimating response time distributions.

We present a statistical model for inference with response time (RT) distributions. The model has the following features. First, it provides a means of estimating the shape, scale, and location (shift) of RT distributions. Second, it is hierarchical and models between-subjects and within-subjects variability simultaneously. Third, inference with the model is Bayesian and provides a principled and efficient means of pooling information across disparate data from different individuals. Because the model efficiently pools information across individuals, it is particularly well suited for those common cases in which the researcher collects a limited number of observations from several participants. Monte Carlo simulations reveal that the hierarchical Bayesian model provides more accurate estimates than several popular competitors do. We illustrate the model by providing an analysis of the symbolic distance effect in which participants can more quickly ascertain the relationship between nonadjacent digits than that between adjacent digits.

Cognition↗

A Beauveria phylogeny inferred from nuclear ITS and EF1-alpha sequences: evidence for cryptic diversification and links to Cordyceps teleomorphs.

Beauveria is a globally distributed genus of soil-borne entomopathogenic hyphomycetes of interest as a model system for the study of entomopathogenesis and the biological control of pest insects. Species recognition in Beauveria is difficult due to a lack of taxonomically informative morphology. This has impeded assessment of species diversity in this genus and investigation of their natural history. A gene-genealogical approach was used to investigate molecular phylogenetic diversity of Beauveria and several presumptively related Cordyceps species. Analyses were based on nuclear ribosomal internal transcribed spacer (ITS) and elongation factor 1-alpha (EF1-alpha) sequences for 86 exemplar isolates from diverse geographic origins, habitats and insect hosts. Phylogenetic trees were inferred using maximum parsimony and Bayesian likelihood methods. Six well supported clades within Beauveria, provisionally designated A-F, were resolved in the EF1-alpha and combined gene phylogenies. Beauveria bassiana, a ubiquitous species that is characterized morphologically by globose to subglobose conidia, was determined to be non-monophyletic and consists of two unrelated lineages, clades A and C. Clade A is globally distributed and includes the Asian teleomorph Cordyceps staphylinidaecola and its probable synonym C. bassiana. All isolates contained in Clade C are anamorphic and originate from Europe and North America. Clade B includes isolates of B. brongniartii, a Eurasian species complex characterized by ellipsoidal conidia. Clade D includes B. caledonica and B. vermiconia, which produce cylindrical and comma-shaped conidia, respectively. Clade E, from Asia, includes Beauveria anamorphs and a Cordyceps teleomorph that both produce ellipsoidal conidia. Clade F, the basal branch in the Beauveria phylogeny includes the South American species B. amorpha, which produces cylindrical conidia. Lineage diversity detected within clades A, B and C suggests that prevailing morphological species concepts underestimate species diversity within these groups. Continental endemism of lineages in B. bassiana s.l. (clades A and C) indicates that isolation by distance has been an important factor in the evolutionary diversification of these clades. Permutation tests indicate that host association is essentially random in both B. bassiana s.l. clades A and C, supporting past assumptions that this species is not host specific. In contrast, isolates in clades B and D occurred primarily on coleopteran hosts, although sampling in these clades was insufficient to assess host affliation at lower taxonomic ranks. The phylogenetic placement of Cordyceps staphylinidaecola/bassiana, and C. scarabaeicola within Beauveria corroborates prior reports of these anamorph-teleomorph connections. These results establish a phylogenetic framework for further taxonomic, phylogenetic and comparative biological investigations of Beauveria and their corresponding Cordyceps teleomorphs.

Animals↗

Evaluating the quality of a probabilistic diagnostic system using different inferencing strategies.

In this paper we describe the evaluation of a probabilistic diagnostic system for patients with renal mass. Three inference models: Multi-membership Bayesian (MB), Minimal Diagnosis (MD) and Bayesian Network (BN), and 72 patients are used to illustrate three interrelated measures of system performance: accuracy, reliability and discriminating power. The inferencing strategies we tested demonstrated the kind of trade-offs in the performance measures that can be expected from imperfect systems. Ultimately, the purpose and expected use of a system should dictate the relative importance ascribed to different aspects of system performance.

Adolescent↗

Molecular systematics of the Eastern Fence Lizard (Sceloporus undulatus): a comparison of Parsimony, Likelihood, and Bayesian approaches.

Phylogenetic analysis of large datasets using complex nucleotide substitution models under a maximum likelihood framework can be computationally infeasible, especially when attempting to infer confidence values by way of nonparametric bootstrapping. Recent developments in phylogenetics suggest the computational burden can be reduced by using Bayesian methods of phylogenetic inference. However, few empirical phylogenetic studies exist that explore the efficiency of Bayesian analysis of large datasets. To this end, we conducted an extensive phylogenetic analysis of the wide-ranging and geographically variable Eastern Fence Lizard (Sceloporus undulatus). Maximum parsimony, maximum likelihood, and Bayesian phylogenetic analyses were performed on a combined mitochondrial DNA dataset (12S and 16S rRNA, ND1 protein-coding gene, and associated tRNA; 3,688 bp total) for 56 populations of S. undulatus (78 total terminals including other S. undulatus group species and outgroups). Maximum parsimony analysis resulted in numerous equally parsimonious trees (82,646 from equally weighted parsimony and 335 from weighted parsimony). The majority rule consensus tree derived from the Bayesian analysis was topologically identical to the single best phylogeny inferred from the maximum likelihood analysis, but required approximately 80% less computational time. The mtDNA data provide strong support for the monophyly of the S. undulatus group and the paraphyly of "S. undulatus" with respect to S. belli, S. cautus, and S. woodi. Parallel evolution of ecomorphs within "S. undulatus" has masked the actual number of species within this group. This evidence, along with convincing patterns of phylogeographic differentiation suggests "S. undulatus" represents at least four lineages that should be recognized as evolutionary species.

Animals↗

Variations over the message computation algorithm of lazy propagation.

Improving the performance of belief updating becomes increasingly important as real-world Bayesian networks continue to grow larger and more complex. In this paper, an investigation is done on how variations over the message-computation algorithm of lazy propagation may impact its performance. Lazy propagation is a junction-tree-based inference algorithm for belief updating in Bayesian networks. Lazy propagation combines variable elimination (VE) with a Shenoy-Shafer message-passing scheme in an attempt to exploit the independence properties induced by evidence in a junction-tree-based algorithm. The authors investigate, the use of arc reversal (AR) and symbolic probabilistic inference (SPI) as alternative algorithms for computing clique-to-clique messages in lazy propagation. The paper presents the results of an empirical evaluation of the performance of lazy propagation using AR, SPI, and VE as the message-computation algorithm. The results of the empirical evaluation show that no single algorithm outperforms or is outperformed by the other two alternatives. In many cases, there is no significant difference in the performance of the three algorithms.

Algorithms↗

Bayesian estimation of ancestral character states on phylogenies.

Biologists frequently attempt to infer the character states at ancestral nodes of a phylogeny from the distribution of traits observed in contemporary organisms. Because phylogenies are normally inferences from data, it is desirable to account for the uncertainty in estimates of the tree and its branch lengths when making inferences about ancestral states or other comparative parameters. Here we present a general Bayesian approach for testing comparative hypotheses across statistically justified samples of phylogenies, focusing on the specific issue of reconstructing ancestral states. The method uses Markov chain Monte Carlo techniques for sampling phylogenetic trees and for investigating the parameters of a statistical model of trait evolution. We describe how to combine information about the uncertainty of the phylogeny with uncertainty in the estimate of the ancestral state. Our approach does not constrain the sample of trees only to those that contain the ancestral node or nodes of interest, and we show how to reconstruct ancestral states of uncertain nodes using a most-recent-common-ancestor approach. We illustrate the methods with data on ribonuclease evolution in the Artiodactyla. Software implementing the methods (BayesMultiState) is available from the authors.

Animals↗

Bayesian techniques for sample size determination in clinical trials: a short review.

The aim of this paper is to review some key techniques of Bayesian methods of sample size determination. The approach is to cover a small number of simple problems, such as estimating the mean of a normal distribution. The methods considered are in two groups: inferential and decision theoretic. In the inferential Bayesian methods of sample size determination, we are solely concerned with the inference about the parameter(s) of interest. The fully Bayesian or decision theoretic approach treats the problem as a decision problem and employs a loss or utility function.

Bayes Theorem↗

Bayesian Mendelian randomization reveals a protective effect of later age at first sexual intercourse against erectile dysfunction.

Erectile dysfunction (ED) is a prevalent health condition with significant psychosocial impacts, yet the causal role of age at first sexual intercourse (AFS) remains unclear. This study investigated the causal effect of AFS on the risk of ED using Mendelian randomization (MR) and Bayesian methods. Five traditional 2-sample MR analyses and 5 Bayesian MR analyses were performed using genome-wide association studies summary statistics from European populations. Sensitivity analyses included MR Egger regression, MR-pleiotropy residual sum and outlier, and Cochran Q-test. In mixed-sex cohorts (Groups 1 and 2), inverse variance weighted results demonstrated significant protective effects: odds ratio (OR) = 0.626, θ = -0.469, P = 2.73 × 10-6 for Group 1 and OR = 0.617, θ = -0.483, P = 3.56 × 10-5 for Group 2. The analyses for male-specific cohorts (Groups 3-10) showed weaker but consistent effects. For Group 3, OR = 0.643, θ = -0.442, P = .010. For Group 4, some instrumental variables associated with confounders were removed. The result became statistically insignificant: OR = 0.680, θ = -0.385, P = .064. For Group 5, the instrument selection criteria were relaxed and significance was retained: OR = 0.695, θ = -0.364, P = .016. For Groups 6 to 10, Bayesian MR was used to strengthen the inferences. In particular, for Group 8, which has a strongly informed prior, a posterior mean θ = -0.358 and a 95% credible interval (-0.575, -0.136) were obtained. This study provides evidence supporting a causal protective effect of later AFS on ED risk. While traditional MR analyses in male-specific cohorts yielded suggestive results, Bayesian MR analyses, which allow for the integration of prior evidence, provided more precise estimates and strengthened the causal inference. These findings may inform future sexual health policies. Strengths include the use of male-specific cohorts and Bayesian enhancement for weak instruments. Limitations include reliance on European-ancestry data and inability to stratify ED subtypes.

Male↗

Inference of Gene Flow between Species from Genomic Data When the Mode, Direction, and Lineages are Misspecified.

Thanks to genomic data, interspecific gene flow is increasingly recognized as a major evolutionary force that shapes biodiversity. Two models have been developed in the multispecies coalescent (MSC) framework to infer gene flow from genomic data, assuming either constant-rate continuous migration (MSC-M) or discrete introgression/hybridization (MSC-I). The extreme simplicity of these models raises concerns about their usefulness as they represent misspecified models when applied to real data. Here, we study inference of gene flow under the MSC-M model, considering mis-assignment of gene flow onto incorrect parental or daughter lineages, misspecification of the direction of gene flow, and misspecification of the mode of gene flow. Mis-assignment of gene flow to an incorrect lineage causes large biases in the estimated rates. The Bayesian test has high power for inferring both recent and ancient gene flow, between either sister lineages or nonsister lineages, although misspecification of the direction of gene flow may make it hard to distinguish early divergence with gene flow from recent complete isolation. Misspecification of the mode of gene flow (MSC-I versus MSC-M) has small local effects, and gene flow is detected with high power despite the misspecification. We analyze a genomic dataset from the purple cone spruce (Picea spp., Pinaceae), which putatively arose through homoploid hybrid speciation, to demonstrate practical implications of our theoretical analyses. Overall, we find that the extremely idealized models of gene flow (in particular the discrete MSC-I model) are very effective for extracting information about species divergence and gene flow from genomic data.

Gene Flow↗

Genetic structure and assignment tests demonstrate illegal translocation of red deer (Cervus elaphus) into a continuous population.

Molecular forensic methods are being increasingly used to help enforce wildlife conservation laws. Using multilocus genotyping, illegal translocation of an animal can be demonstrated by excluding all potential source populations as an individual's population of origin. Here, we illustrate how this approach can be applied to a large continuous population by defining the population genetic structure and excluding suspect animals from each identified cluster. We aimed to test the hypothesis that recreational hunters had illegally introduced a group of red deer into a hunting area in Luxembourg. Reference samples were collected over a large area in order to test the possibility that the suspect individuals might be recent immigrants. Due to isolation-by-distance relationships in the data set, inferring the number of genetic clusters using Bayesian methods was not straightforward. Biologically meaningful clusters were only obtained by simultaneously analysing spatial and genetic information using the program baps 4.1. We inferred the presence of three genetic clusters in the study region. Using partial Mantel tests, we detected barriers to gene flow other than distance, probably created by a combination of urban areas, motorways and a river valley used for viticulture. The four focal animals could be excluded with a high certainty from the three genetic subpopulations and it was therefore likely that they had been released illegally.

Animal Migration↗

A Bayesian compound stochastic process for modeling nonstationary and nonhomogeneous sequence evolution.

Variations of nucleotidic composition affect phylogenetic inference conducted under stationary models of evolution. In particular, they may cause unrelated taxa sharing similar base composition to be grouped together in the resulting phylogeny. To address this problem, we developed a nonstationary and nonhomogeneous model accounting for compositional biases. Unlike previous nonstationary models, which are branchwise, that is, assume that base composition only changes at the nodes of the tree, in our model, the process of compositional drift is totally uncoupled from the speciation events. In addition, the total number of events of compositional drift distributed across the tree is directly inferred from the data. We implemented the method in a Bayesian framework, relying on Markov Chain Monte Carlo algorithms, and applied it to several nucleotidic data sets. In most cases, the stationarity assumption was rejected in favor of our nonstationary model. In addition, we show that our method is able to resolve a well-known artifact. By Bayes factor evaluation, we compared our model with 2 previously developed nonstationary models. We show that the coupling between speciations and compositional shifts inherent to branchwise models may lead to an overparameterization, resulting in a lesser fit. In some cases, this leads to incorrect conclusions, concerning the nature of the compositional biases. In contrast, our compound model more flexibly adapts its effective number of parameters to the data sets under investigation. Altogether, our results show that accounting for nonstationary sequence evolution may require more elaborate and more flexible models than those currently used.

Animals↗

Predicting the effect of missense mutations on protein function: analysis with Bayesian networks.

BACKGROUND: A number of methods that use both protein structural and evolutionary information are available to predict the functional consequences of missense mutations. However, many of these methods break down if either one of the two types of data are missing. Furthermore, there is a lack of rigorous assessment of how important the different factors are to prediction. RESULTS: Here we use Bayesian networks to predict whether or not a missense mutation will affect the function of the protein. Bayesian networks provide a concise representation for inferring models from data, and are known to generalise well to new data. More importantly, they can handle the noisy, incomplete and uncertain nature of biological data. Our Bayesian network achieved comparable performance with previous machine learning methods. The predictive performance of learned model structures was no better than a naïve Bayes classifier. However, analysis of the posterior distribution of model structures allows biologically meaningful interpretation of relationships between the input variables. CONCLUSION: The ability of the Bayesian network to make predictions when only structural or evolutionary data was observed allowed us to conclude that structural information is a significantly better predictor of the functional consequences of a missense mutation than evolutionary information, for the dataset used. Analysis of the posterior distribution of model structures revealed that the top three strongest connections with the class node all involved structural nodes. With this in mind, we derived a simplified Bayesian network that used just these three structural descriptors, with comparable performance to that of an all node network.

Algorithms↗

Bayesian modeling of multiple lesion onset and growth from interval-censored data.

In studying rates of occurrence and progression of lesions (or tumors), it is typically not possible to obtain exact onset times for each lesion. Instead, data consist of the number of lesions that reach a detectable size between screening examinations, along with measures of the size/severity of individual lesions at each exam time. This interval-censored data structure makes it difficult to properly adjust for the onset time distribution in assessing covariate effects on rates of lesion progression. This article proposes a joint model for the multiple lesion onset and progression process, motivated by cross-sectional data from a study of uterine leiomyoma tumors. By using a joint model, one can potentially obtain more precise inferences on rates of onset, while also performing onset time-adjusted inferences on lesion severity. Following a Bayesian approach, we propose a data augmentation Markov chain Monte Carlo algorithm for posterior computation.

Adult↗

PSMIX: an R package for population structure inference via maximum likelihood method.

BACKGROUND: Inference of population stratification and individual admixture from genetic markers is an integrative part of a study in diverse situations, such as association mapping and evolutionary studies. Bayesian methods have been proposed for population stratification and admixture inference using multilocus genotypes and widely used in practice. However, these Bayesian methods demand intensive computation resources and may run into convergence problem in Markov Chain Monte Carlo based posterior samplings. RESULTS: We have developed PSMIX, an R package based on maximum likelihood method using expectation-maximization algorithm, for inference of population stratification and individual admixture. CONCLUSION: Compared with software based on Bayesian methods (e.g., STRUCTURE), PSMIX has similar accuracy, but more efficient computations.PSMIX and its supplemental documents are freely available at http://bioinformatics.med.yale.edu/PSMIX.

Algorithms↗

Model-based multi-locus estimation of decapod phylogeny and divergence times.

Phylogenetic relationships among all of the major decapod infraorders have never been estimated using molecular data, while morphological studies produce conflicting results. In the present study, the phylogenetic relationships among the decapod basal suborder Dendrobranchiata and all of the currently recognized decapod infraorders within the suborder Pleocyemata (Caridea, Stenopodidea, Achelata, Astacidea, Thalassinidea, Anomala, and Brachyura) were inferred using 16S mtDNA, 18S and 28S rRNA, and the histone H3 gene. Phylogenies were reconstructed using the model-based methods of maximum likelihood and Bayesian methods coupled with Markov Chain Monte Carlo inference. The phylogenies revealed that the seven infraorders are monophyletic, with high clade support values (bp>70; pP>0.95) under both methods. The two suborders also were recovered as monophyletic, but with weaker support (bp=70; pP=0.74). Although the nodal support values for infraordinal relationships were low (bp<50; pP<0.77) the Anomala and Brachyura were basal to the rest of the 'Reptantia' in both reconstructions and using Bayesian tree topology tests alternate morphology-based hypotheses were rejected (P<0.01). Newly developed multi-locus Bayesian and likelihood heuristic rate-smoothing methods to estimate divergence times were compared using eight fossil and geological calibrations. Estimated times revealed that the Decapoda originated earlier than 437MYA and that the radiation within the group occurred rapidly, with all of the major lineages present by 325MYA. Node time estimation under both approaches is severely affected by the number and phylogenetic distribution of the fossil calibrations chosen. For analyses incorporating fossils as fixed ages, more consistent results were obtained by using both shallow and deep or clade-related calibration points. Divergence time estimation using fossils as lower and upper limits performed well with as few as one upper limit and a single deep fossil lower limit calibration.

Animals↗