Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

MRDtarget: A heuristic Gaussian approach for optimizing targeted capture regions to enhance Minimal Residual Disease detection.

Molecular residual disease (MRD) detection, initially developed for hematologic malignancies, has become a critical biomarker for monitoring solid tumors. MRD detection primarily relies on circulating tumor DNA (ctDNA) analysis using next-generation sequencing, offering high sensitivity and broad genomic coverage. However, challenges remain in designing cost-effective panels that maximize mutation detection while maintaining biological relevance. Fixed panels often lack sufficient patient-specific mutation coverage, while WES-based personalized MRD assays, despite their high sensitivity, are costly and less accessible. We developed a tumor comprehensive genomic profiling (CGP)-informed personalized MRD assay to detect tumor-derived mutations, which allowed us to design patient-specific personalized panels and meanwhile, provide a cost-effective alternative to whole exome sequencing (WES). To address these limitations, we developed MRDtarget, a heuristic multivariate Gaussian model-based targeted capture region selection method. By expanding beyond traditional hotspot regions, MRDtarget optimizes variant tracking for MRD detection, significantly improving sensitivity. Using a Bayesian inference-based heuristic approach, MRDtarget integrates multi-feature informativeness rates to identify optimal genomic regions for capture. Experimental results demonstrate that MRDtarget enables the detection of more variants per patient. This study underscores the importance of rational panel design to improve MRD sensitivity and provides a novel approach to enhance precision diagnostics and treatment for solid tumor patients.

Humans↗

Using temporally spaced sequences to simultaneously estimate migration rates, mutation rate and population sizes in measurably evolving populations.

We present a Bayesian statistical inference approach for simultaneously estimating mutation rate, population sizes, and migration rates in an island-structured population, using temporal and spatial sequence data. Markov chain Monte Carlo is used to collect samples from the posterior probability distribution. We demonstrate that this chain implementation successfully reaches equilibrium and recovers truth for simulated data. A real HIV DNA sequence data set with two demes, semen and blood, is used as an example to demonstrate the method by fitting asymmetric migration rates and different population sizes. This data set exhibits a bimodal joint posterior distribution, with modes favoring different preferred migration directions. This full data set was subsequently split temporally for further analysis. Qualitative behavior of one subset was similar to the bimodal distribution observed with the full data set. The temporally split data showed significant differences in the posterior distributions and estimates of parameter values over time.

Biological Evolution↗

Divergent selection for uterine capacity in rabbits. I. Genetic parameters and response to selection.

A 10-generation divergent selection experiment for uterine capacity (UC) measured as litter size in unilaterally ovariectomized females was carried out in rabbits. A total of 2,996 observations on uterine capacity of does (up to four parities) was recorded. Laparoscopy was performed at d 12 of their second gestation, and ovulation rate (OR) and number of implanted embryos (IE) were recorded in 735 does. Prenatal survival (PS) was assessed as UC/OR, embryo survival (ES) as IE/OR, and fetal survival (FS) as UC/IE. Genetic parameters and genetic trends were inferred using Bayesian methods. Marginal posterior distributions of all unknowns were estimated by Gibbs sampling. Heritabilities of UC, OR, IE, ES, FS, and PS were 0.11, 0.32, 0.22, 0.04, 0.14, and 0.09, respectively. Genetic and phenotypic correlations between FS and ES were low, suggesting different biological mechanisms for the two periods of survival. After 10 generations of selection, the divergence was approximately 1.5 rabbits, or approximately 1% per generation. Approximately one-half of this response was obtained in the first two generations of selection, which may suggest the presence of a major gene segregating in the base population.

Animals↗

Divergent selection for uterine capacity in rabbits. II. Correlated response in litter size and its components estimated with a cryopreserved control population.

Our objective was to evaluate the correlated responses to selection for litter size and its components after 10 generations of divergent selection for uterine capacity (UC). A total of 294 intact females from the 11th and 12th generations of divergent selection for high and low UC and from a cryopreserved control population was used (139, 112, and 43 females, respectively). Uterine capacity was assessed as litter size in unilaterally ovariectomized females. Traits recorded on females for up to five parities were litter size (LS) and number born alive (NBA). Laparoscopy was performed in all females at d 12 of their second parity, and the ovulation rate (OR) and number of implanted embryos (IE) were recorded in these females. Embryo survival (ES = IE/OR), fetal survival (FS = LS/IE), and prenatal survival (PS = LS/OR) were computed. Correlated responses in LS and in its components were inferred using Bayesian methods. Correlated responses in LS were asymmetric. The divergence between high and low lines was 2.35 kits, mainly because of a higher correlated response in the low line (1.88 kits). The lower LS in the low line was associated with a lower PS (control - low = 0.14), because of decreases in ES and FS.

Animals↗

Informative structure priors: joint learning of dynamic regulatory networks from multiple types of data.

We present a method for jointly learning dynamic models of transcriptional regulatory networks from gene expression data and transcription factor binding location data. Models are automatically learned using dynamic Bayesian network inference algorithms; joint learning is accomplished by incorporating evidence from gene expression data through the likelihood, and from transcription factor binding location data through the prior. We propose a new informative structure prior with two advantages. First, the prior incorporates evidence from location data probabilistically, allowing it to be weighed against evidence from expression data. Second, the prior takes on a factorable form that is computationally efficient when learning dynamic regulatory networks. Results obtained from both simulated and experimental data from the yeast cell cycle demonstrate that this joint learning algorithm can recover dynamic regulatory networks from multiple types of data that are more accurate than those recovered from each type of data in isolation.

Bayes Theorem↗

An empirical examination of the standard errors of maximum likelihood phylogenetic parameters under the molecular clock via bootstrapping.

The molecular clock theory has greatly enlightened our understanding of macroevolutionary events. Maximum likelihood (ML) estimation of divergence times involves the adoption of fixed calibration points, and the confidence intervals associated with the estimates are generally very narrow. The credibility intervals are inferred assuming that the estimates are normally distributed, which may not be the case. Moreover, calculation of standard errors is usually carried out by the curvature method and is complicated by the difficulty in approximating second derivatives of the likelihood function. In this study, a standard primate phylogeny was used to examine the standard errors of ML estimates via the bootstrap method. Confidence intervals were also assessed from the posterior distribution of divergence times inferred via Bayesian Markov Chain Monte Carlo. For the primate topology under evaluation, no significant differences were found between the bootstrap and the curvature methods. Also, Bayesian confidence intervals were always wider than those obtained by ML.

Animals↗

Computerized expert system for the diagnosis of pulp-related pain.

A major problem in the correct diagnosis of pulpal pain is that the associated clinical signs do not predictably correlate with the underlying pathological process. Using conditional probabilities of various pulp conditions from published data, Bayesian Statistical Inference provides the means for deriving a composite probability of the presence of a disease from a multiple set of symptoms. A computer program that can infer a diagnosis for pulpal pain from any combination of 17 clinical symptoms has been developed. From the data, the program provides the computed relative probabilities of a healthy pulp, a saveable pulp, an unsaveable pulp, and a necrotic pulp being present.

Bayes Theorem↗

[The Bayesian approach: another way of drawing inferences].

The serious objections made with regard to significance tests account for the necessity of employing another inferential procedure. Bayesian methods are free of such objections, and they provide a very attractive alternative. By means of a simple example, this article illustrates how a typical problem of medical research could be solved using these two approaches. Bayesian methods offer more information and are more useful than conventional ones for analyzing experimental outcomes. In addition, a natural interpretation of conclusions are given by Bayesian methods. Finally, modern computational programs allow us to solve their complex calculation.

Bayes Theorem↗

Comparative evaluation of reverse engineering gene regulatory networks with relevance networks, graphical gaussian models and bayesian networks.

MOTIVATION: An important problem in systems biology is the inference of biochemical pathways and regulatory networks from postgenomic data. Various reverse engineering methods have been proposed in the literature, and it is important to understand their relative merits and shortcomings. In the present paper, we compare the accuracy of reconstructing gene regulatory networks with three different modelling and inference paradigms: (1) Relevance networks (RNs): pairwise association scores independent of the remaining network; (2) graphical Gaussian models (GGMs): undirected graphical models with constraint-based inference, and (3) Bayesian networks (BNs): directed graphical models with score-based inference. The evaluation is carried out on the Raf pathway, a cellular signalling network describing the interaction of 11 phosphorylated proteins and phospholipids in human immune system cells. We use both laboratory data from cytometry experiments as well as data simulated from the gold-standard network. We also compare passive observations with active interventions. RESULTS: On Gaussian observational data, BNs and GGMs were found to outperform RNs. The difference in performance was not significant for the non-linear simulated data and the cytoflow data, though. Also, we did not observe a significant difference between BNs and GGMs on observational data in general. However, for interventional data, BNs outperform GGMs and RNs, especially when taking the edge directions rather than just the skeletons of the graphs into account. This suggests that the higher computational costs of inference with BNs over GGMs and RNs are not justified when using only passive observations, but that active interventions in the form of gene knockouts and over-expressions are required to exploit the full potential of BNs. AVAILABILITY: Data, software and supplementary material are available from http://www.bioss.sari.ac.uk/staff/adriano/research.html

Algorithms↗

Comparison of site-specific rate-inference methods for protein sequences: empirical Bayesian methods are superior.

The degree to which an amino acid site is free to vary is strongly dependent on its structural and functional importance. An amino acid that plays an essential role is unlikely to change over evolutionary time. Hence, the evolutionary rate at an amino acid site is indicative of how conserved this site is and, in turn, allows evaluation of its importance in maintaining the structure/function of the protein. When using probabilistic methods for site-specific rate inference, few alternatives are possible. In this study we use simulations to compare the maximum-likelihood and Bayesian paradigms. We study the dependence of inference accuracy on such parameters as number of sequences, branch lengths, the shape of the rate distribution, and sequence length. We also study the possibility of simultaneously estimating branch lengths and site-specific rates. Our results show that a Bayesian approach is superior to maximum-likelihood under a wide range of conditions, indicating that the prior that is incorporated into the Bayesian computation significantly improves performance. We show that when branch lengths are unknown, it is better first to estimate branch lengths and then to estimate site-specific rates. This procedure was found to be superior to estimating both the branch lengths and site-specific rates simultaneously. Finally, we illustrate the difference between maximum-likelihood and Bayesian methods when analyzing site-conservation for the apoptosis regulator protein Bcl-x(L).

Animals↗

Improved statistical inference from DNA microarray data using analysis of variance and a Bayesian statistical framework. Analysis of global gene expression in Escherichia coli K12.

We describe statistical methods based on the t test that can be conveniently used on high density array data to test for statistically significant differences between treatments. These t tests employ either the observed variance among replicates within treatments or a Bayesian estimate of the variance among replicates within treatments based on a prior estimate obtained from a local estimate of the standard deviation. The Bayesian prior allows statistical inference to be made from microarray data even when experiments are only replicated at nominal levels. We apply these new statistical tests to a data set that examined differential gene expression patterns in IHF(+) and IHF(-) Escherichia coli cells (Arfin, S. M., Long, A. D., Ito, E. T., Tolleri, L., Riehle, M. M., Paegle, E. S., and Hatfield, G. W. (2000) J. Biol. Chem. 275, 29672-29684). These analyses identify a more biologically reasonable set of candidate genes than those identified using statistical tests not incorporating a Bayesian prior. We also show that statistical tests based on analysis of variance and a Bayesian prior identify genes that are up- or down-regulated following an experimental manipulation more reliably than approaches based only on a t test or fold change. All the described tests are implemented in a simple-to-use web interface called Cyber-T that is located on the University of California at Irvine genomics web site.

Bayes Theorem↗

Bayesian sample-size determination for inference on two binomial populations with no gold standard classifier.

We consider the impact of test properties on the required sample size for the Bayesian design problem for comparing two proportions with error-prone data. Specifically, we examine four cases: a single diagnostic test and two independent diagnostic tests, both when the test properties are identical across populations and when they differ. Interval-based and moment-based sample-size determination criteria are contrasted using Monte Carlo simulation methods. We consider an application in which Strongyloides infections are compared in two populations.

Animals↗

Sensitivity and specificity of inferring genetic regulatory interactions from microarray experiments with dynamic Bayesian networks.

MOTIVATION: Bayesian networks have been applied to infer genetic regulatory interactions from microarray gene expression data. This inference problem is particularly hard in that interactions between hundreds of genes have to be learned from very small data sets, typically containing only a few dozen time points during a cell cycle. Most previous studies have assessed the inference results on real gene expression data by comparing predicted genetic regulatory interactions with those known from the biological literature. This approach is controversial due to the absence of known gold standards, which renders the estimation of the sensitivity and specificity, that is, the true and (complementary) false detection rate, unreliable and difficult. The objective of the present study is to test the viability of the Bayesian network paradigm in a realistic simulation study. First, gene expression data are simulated from a realistic biological network involving DNAs, mRNAs, inactive protein monomers and active protein dimers. Then, interaction networks are inferred from these data in a reverse engineering approach, using Bayesian networks and Bayesian learning with Markov chain Monte Carlo. RESULTS: The simulation results are presented as receiver operator characteristics curves. This allows estimating the proportion of spurious gene interactions incurred for a specified target proportion of recovered true interactions. The findings demonstrate how the network inference performance varies with the training set size, the degree of inadequacy of prior assumptions, the experimental sampling strategy and the inclusion of further, sequence-based information. AVAILABILITY: The programs and data used in the present study are available from http://www.bioss.sari.ac.uk/~dirk/Supplements

Bayes Theorem↗

Bayesian analysis of population PK/PD models: general concepts and software.

Markov chain Monte Carlo (MCMC) techniques have revolutionized the field of Bayesian statistics by enabling posterior inference for arbitrarily complex models. The now widely used WinBUGS software has, over the years, made the methodology accessible to a great many applied scientists, in all fields of research. Despite this, serious application of MCMC methods within the field of population PK/PD has been comparatively limited. We appreciate that for many applied pharmacokineticists the prospect of conducting a Bayesian analysis will require numerous alien concepts to be taken on board and it may be difficult to justify investing the time and effort required in order to understand them (especially since the approach is so computer-intensive). For this reason we provide here a thorough (but often informal) discussion of all aspects of Bayesian inference as they apply specifically to population PK/PD. We also acknowledge that while the WinBUGS software is general purpose, model specification for some types of problem, population PK/PD being a prime example, can be very difficult, to the extent that a specialized interface for describing the problem at hand is often a practical necessity. In the latter part of this paper we describe such an interface, namely PKBugs. A principal aim of the paper is to offer sufficient technical background, in an easy to follow format, that the reader may develop both the confidence and know-how to make appropriate use of the PKBugs/WinBUGS framework (or similar software) for their own data analysis needs, should they choose to adopt a Bayesian approach.

Bayes Theorem↗

Bayesian hypothesis testing of four-taxon topologies using molecular sequence data.

The reconstruction of phylogenetic trees from molecular sequences presents unusual problems for statistical inference. For example, three possible alternatives must be considered for four taxa when inferring the correct unrooted tree (referred to as a topology). In our view, classical hypothesis testing is poorly suited to this triangular set of alternative hypotheses. In this article, we develop Bayesian inference to determine the posterior probability that a four-taxon topology is correct given the sequence data and the evolutionary parsimony algorithm for phylogenetic reconstruction. We assess the frequency properties of our models in a large simulation study. Bayesian inference under the principles of evolutionary parsimony is shown to be well calibrated with reasonable discriminating power for a wide range of realistic conditions, including conditions that violate the assumptions of evolutionary parsimony.

Base Sequence↗

A Bayesian heterogeneous analysis of variance approach to inferring recent selective sweeps.

The distribution of microsatellite allele sizes in populations aids in understanding the genetic diversity of species and the evolutionary history of recent selective sweeps. We propose a heterogeneous Bayesian analysis of variance model for inferring loci involved in recent selective sweeps by analyzing the distribution of allele sizes at multiple loci in multiple populations. Our model is shown to be consistent with a multilocus test statistic, ln RV, proposed for identifying microsatellite loci involved in recent selective sweeps. Our methodology differs in that it accepts original allele size data rather than summary statistics and allows the incorporation of prior knowledge about allele frequencies using a hierarchical prior distribution consisting of log normal and gamma probability distributions. Interesting features of the model are its ability to simultaneously analyze allele size data for any number of populations and to cope with the presence of any number of selected loci. The utility of the method is illustrated by application to two sets of microsatellite allele size data for a group of West African Anopheles gambiae populations. The results are consistent with the suppressed-recombination model of speciation, and additional candidate loci on chromosomes 2 (079 and 175) and 3 (088) are discovered that escaped former analysis.

Alleles↗

Inference for the cost-effectiveness acceptability curve and cost-effectiveness ratio.

The aim of this article is to consider Bayesian and frequentist inference methods for measures of incremental cost effectiveness in data obtained via a clinical trial. The most useful measure is the cost-effectiveness (C/E) acceptability curve. Recent publications on Bayesian estimation have assumed a normal posterior distribution, which ignores uncertainty in estimated variances, and suggest unnecessarily complicated methods of computation. We present a simple Bayesian computation for the C/E acceptability curve and a simple frequentist analogue. Our approach takes account of errors in estimated variances, resulting in calculations that are based on distributions rather than normal distributions. If inference is required about the C/E ratio, we argue that the standard frequentist procedures give unreliable or misleading inferences, and present instead a Bayesian interval.

Antirheumatic Agents↗

Molecular systematics of the Jacks (Perciformes: Carangidae) based on mitochondrial cytochrome b sequences using parsimony, likelihood, and Bayesian approaches.

The Carangidae represent a diverse family of marine fishes that include both ecologically and economically important species. Currently, there are four recognized tribes within the family, but phylogenetic relationships among them based on morphology are not resolved. In addition, the tribe Carangini contains species with a variety of body forms and no study has tried to interpret the evolution of this diversity. We used DNA sequences from the mitochondrial cytochrome b gene to reconstruct the phylogenetic history of 50 species from each of the four tribes of Carangidae and four carangoid outgroup taxa. We found support for the monophyly of three tribes within the Carangidae (Carangini, Naucratini, and Trachinotini); however, monophyly of the fourth tribe (Scomberoidini) remains questionable. A sister group relationship between the Carangini and the Naucratini is well supported. This clade is apparently sister to the Trachinotini plus Scomberoidini but there is uncertain support for this relationship. Additionally, we examined the evolution of body form within the tribe Carangini and determined that each of the predominant clades has a distinct evolutionary trend in body form. We tested three methods of phylogenetic inference, parsimony, maximum-likelihood, and Bayesian inference. Whereas the three analyses produced largely congruent hypotheses, they differed in several important relationships. Maximum-likelihood and Bayesian methods produced hypotheses with higher support values for deep branches. The Bayesian analysis was computationally much faster and yet produced phylogenetic hypotheses that were very similar to those of the maximum-likelihood analysis.

Animals↗