Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Inferring gene transcriptional modulatory relations: a genetical genomics approach.

Bayesian network modeling is a promising approach to define and evaluate gene expression circuits in diverse tissues and cell types under different experimental conditions. The power and practicality of this approach can be improved by restricting the number of potential interactions among genes and by defining causal relations before evaluating posterior probabilities for billions of networks. A newly developed genetical genomics method that combines transcriptome profiling with complex trait analysis now provides strong constraints on network architecture. This method detects those chromosomal intervals responsible for differences in mRNA expression using quantitative trait locus (QTL) mapping. We have developed an efficient Bayesian approach that exploits the genetical genomics method to focus computational effort on the most plausible gene modulatory networks. We exploit a dense marker map for a genetic reference population (GRP) that consists of 32 BXD strains of mice made by intercrossing two progenitor strains--C57BL/6J and DBA/2J. These progenitors differ at approximately 1.3 million known single nucleotide polymorphisms (SNPs), all of which can be exploited to estimate the probability that a gene contains functional polymorphisms that segregate within the GRP. We constructed 66 candidate networks that include all the candidate modulator genes located in the 209 statistically significant trans-acting QTL regions. SNPs that distinguish between the two progenitor strains were used to further winnow the list of candidate modulators. Bayesian network was then used to identify the genetic modulatory relations that best explain the microarray data.

Algorithms↗

Protein molecular function prediction by Bayesian phylogenomics.

We present a statistical graphical model to infer specific molecular function for unannotated protein sequences using homology. Based on phylogenomic principles, SIFTER (Statistical Inference of Function Through Evolutionary Relationships) accurately predicts molecular function for members of a protein family given a reconciled phylogeny and available function annotations, even when the data are sparse or noisy. Our method produced specific and consistent molecular function predictions across 100 Pfam families in comparison to the Gene Ontology annotation database, BLAST, GOtcha, and Orthostrapper. We performed a more detailed exploration of functional predictions on the adenosine-5'-monophosphate/adenosine deaminase family and the lactate/malate dehydrogenase family, in the former case comparing the predictions against a gold standard set of published functional characterizations. Given function annotations for 3% of the proteins in the deaminase family, SIFTER achieves 96% accuracy in predicting molecular function for experimentally characterized proteins as reported in the literature. The accuracy of SIFTER on this dataset is a significant improvement over other currently available methods such as BLAST (75%), GeneQuiz (64%), GOtcha (89%), and Orthostrapper (11%). We also experimentally characterized the adenosine deaminase from Plasmodium falciparum, confirming SIFTER's prediction. The results illustrate the predictive power of exploiting a statistical model of function evolution in phylogenomic problems. A software implementation of SIFTER is available from the authors.

Adenosine Deaminase↗

Assessing the accuracy of ancestral protein reconstruction methods.

The phylogenetic inference of ancestral protein sequences is a powerful technique for the study of molecular evolution, but any conclusions drawn from such studies are only as good as the accuracy of the reconstruction method. Every inference method leads to errors in the ancestral protein sequence, resulting in potentially misleading estimates of the ancestral protein's properties. To assess the accuracy of ancestral protein reconstruction methods, we performed computational population evolution simulations featuring near-neutral evolution under purifying selection, speciation, and divergence using an off-lattice protein model where fitness depends on the ability to be stable in a specified target structure. We were thus able to compare the thermodynamic properties of the true ancestral sequences with the properties of "ancestral sequences" inferred by maximum parsimony, maximum likelihood, and Bayesian methods. Surprisingly, we found that methods such as maximum parsimony and maximum likelihood that reconstruct a "best guess" amino acid at each position overestimate thermostability, while a Bayesian method that sometimes chooses less-probable residues from the posterior probability distribution does not. Maximum likelihood and maximum parsimony apparently tend to eliminate variants at a position that are slightly detrimental to structural stability simply because such detrimental variants are less frequent. Other properties of ancestral proteins might be similarly overestimated. This suggests that ancestral reconstruction studies require greater care to come to credible conclusions regarding functional evolution. Inferred functional patterns that mimic reconstruction bias should be reevaluated.

Algorithms↗

Model selection for mixtures of mutagenetic trees.

The evolution of drug resistance in HIV is characterized by the accumulation of resistance-associated mutations in the HIV genome. Mutagenetic trees, a family of restricted Bayesian tree models, have been applied to infer the order and rate of occurrence of these mutations. Understanding and predicting this evolutionary process is an important prerequisite for the rational design of antiretroviral therapies. In practice, mixtures models of K mutagenetic trees provide more flexibility and are often more appropriate for modelling observed mutational patterns. Here, we investigate the model selection problem for K-mutagenetic trees mixture models. We evaluate several classical model selection criteria including cross-validation, the Bayesian Information Criterion (BIC), and the Akaike Information Criterion. We also use the empirical Bayes method by constructing a prior probability distribution for the parameters of a mutagenetic trees mixture model and deriving the posterior probability of the model. In addition to the model dimension, we consider the redundancy of a mixture model, which is measured by comparing the topologies of trees within a mixture model. Based on the redundancy, we propose a new model selection criterion, which is a modification of the BIC. Experimental results on simulated and on real HIV data show that the classical criteria tend to select models with far too many tree components. Only cross-validation and the modified BIC recover the correct number of trees and the tree topologies most of the time. At the same optimal performance, the runtime of the new BIC modification is about one order of magnitude lower. Thus, this model selection criterion can also be used for large data sets for which cross-validation becomes computationally infeasible.

Bayes Theorem↗

Bayesian analysis of experimental epidemics of foot-and-mouth disease.

We investigate the transmission dynamics of a certain type of foot-and-mouth disease (FMD) virus under experimental conditions. Previous analyses of experimental data from FMD outbreaks in non-homogeneously mixing populations of sheep have suggested a decline in viraemic level through serial passage of the virus, but these do not take into account possible variation in the length of the chain of viral transmission for each animal, which is implicit in the non-observed transmission process. We consider a susceptible-exposed-infectious-removed non-Markovian compartmental model for partially observed epidemic processes, and we employ powerful methodology (Markov chain Monte Carlo) for statistical inference, to address epidemiological issues under a Bayesian framework that accounts for all available information and associated uncertainty in a coherent approach. The analysis allows us to investigate the posterior distribution of the hidden transmission history of the epidemic, and thus to determine the effect of the length of the infection chain on the recorded viraemic levels, based on the posterior distribution of a p-value. Parameter estimates of the epidemiological characteristics of the disease are also obtained. The results reveal a possible decline in viraemia in one of the two experimental outbreaks. Our model also suggests that individual infectivity is related to the level of viraemia.

Animals↗

A probabilistic rule-based expert system.

This paper explores a medical expert system combining techniques of Bayesian network modelling with ideas of weighted inference rules. The weights of the individual rules can be estimated objectively from a training set of actual cases; and they can be used in a Monte Carlo stimulation to estimate objectively conditional probabilities of diagnosis given particular combinations of symptoms. The paper describes and evaluates a medical expert system built according to this design. The diagnostic accuracy of the program was found to be similar to that obtained through the usual application of Bayes theorem with the assumption of conditional independence of symptoms given disease, even though the Bayesian classifier has more than 70 times as many numerical parameters. The method may be promising in cases where small training sets do not permit accurate estimation of large numbers of parameters.

Abdominal Pain↗

A Bayesian mixture model for partitioning gene expression data.

In recent years there has been great interest in making inference for gene expression data collected over time. In this article, we describe a Bayesian hierarchical mixture model for partitioning such data. While conventional approaches cluster the observed data, we assume a nonparametric, random walk model, and partition on the basis of the parameters of this model. The model is flexible and can be tuned to the specific context, respects the order of observations within each curve, acknowledges measurement error, and allows prior knowledge on parameters to be incorporated. The number of partitions may also be treated as unknown, and inferred from the data, in which case computation is carried out via a birth-death Markov chain Monte Carlo algorithm. We first examine the behavior of the model on simulated data, along with a comparison with more conventional approaches, and then analyze meiotic expression data collected over time on fission yeast genes.

Bayes Theorem↗

Frequentist properties of Bayesian posterior probabilities of phylogenetic trees under simple and complex substitution models.

What does the posterior probability of a phylogenetic tree mean?This simulation study shows that Bayesian posterior probabilities have the meaning that is typically ascribed to them; the posterior probability of a tree is the probability that the tree is correct, assuming that the model is correct. At the same time, the Bayesian method can be sensitive to model misspecification, and the sensitivity of the Bayesian method appears to be greater than the sensitivity of the nonparametric bootstrap method (using maximum likelihood to estimate trees). Although the estimates of phylogeny obtained by use of the method of maximum likelihood or the Bayesian method are likely to be similar, the assessment of the uncertainty of inferred trees via either bootstrapping (for maximum likelihood estimates) or posterior probabilities (for Bayesian estimates) is not likely to be the same. We suggest that the Bayesian method be implemented with the most complex models of those currently available, as this should reduce the chance that the method will concentrate too much probability on too few trees.

Animals↗

Causal assessment of Bayesian gene regulatory networks from single-cell transcriptomics.

Gene regulatory network (GRN) inference is an essential tool for revealing dysregulated relationships between genes in different cell types from single-cell transcriptomic (SCT) data. GRNs based on Bayesian networks (BNs) learned from SCT data can elucidate directed regulatory relationships representing complex disease mechanisms and their interplay through graphical modeling. However, software for learning BNs from SCT data is not widely available, nor is software for evaluating the BNs' structural accuracy in representing causal relationships between genes. Here, we describe the scstruc R package. This package provides a suite of BN structure learning algorithms specifically designed to handle SCT data, to evaluate the resulting networks based on the causal relationships they represent regardless of the availability of established molecular interaction networks, and to compare regulatory relationships between conditions. We demonstrated that scstruc can identify biologically relevant differential regulatory relationships between groups on a per-cell basis.

Bayesian networks↗

Genetic parameters for stillbirth in Danish Holstein cows using a Bayesian threshold model.

The objective of this study was to make an inference about the direct and maternal genetic variation of stillbirth for first-calving Holstein cows and to estimate the effect of breed and heterosis for original Danish black and white and Holstein-Friesian. A Bayesian threshold model, which included correlated genetic effects of sires and maternal grandsires was used. Marginal posterior distributions of effects were obtained using Gibbs sampling. Point estimates were compared with results from a linear model using REML. Data with and without twins were analyzed and models with and without effects of breed and heterosis were fitted, but estimates of genetic parameters were almost identical. In all the analyses with threshold models, the marginal posterior mean (and standard deviation) was 0.10 (0.014) for the direct heritability, 0.13 (0.015) for the maternal heritability, and 0.05 (0.10) for the genetic correlation between direct and maternal effects. The stillbirth rate tended to increase with a higher proportion of Holstein-Friesian in the calf and in the dam, but no effects of breed and heterosis were significant. Joint sampling of all location parameters was found superior to univariate sampling in terms of much better mixing properties of the fixed effects. Based on the results showing genetic variation for stillbirth at first calving, both the direct and the maternal effect could be included in the breeding program.

Animals↗

Bayesian model averaging in EEG/MEG imaging.

In this paper, the Bayesian Theory is used to formulate the Inverse Problem (IP) of the EEG/MEG. This formulation offers a comparison framework for the wide range of inverse methods available and allows us to address the problem of model uncertainty that arises when dealing with different solutions for a single data. In this case, each model is defined by the set of assumptions of the inverse method used, as well as by the functional dependence between the data and the Primary Current Density (PCD) inside the brain. The key point is that the Bayesian Theory not only provides for posterior estimates of the parameters of interest (the PCD) for a given model, but also gives the possibility of finding posterior expected utilities unconditional on the models assumed. In the present work, this is achieved by considering a third level of inference that has been systematically omitted by previous Bayesian formulations of the IP. This level is known as Bayesian model averaging (BMA). The new approach is illustrated in the case of considering different anatomical constraints for solving the IP of the EEG in the frequency domain. This methodology allows us to address two of the main problems that affect linear inverse solutions (LIS): (a) the existence of ghost sources and (b) the tendency to underestimate deep activity. Both simulated and real experimental data are used to demonstrate the capabilities of the BMA approach, and some of the results are compared with the solutions obtained using the popular low-resolution electromagnetic tomography (LORETA) and its anatomically constraint version (cLORETA).

Artifacts↗

Bayesian methods for phase I clinical trials.

Phase I clinical trials are conducted to determine the dose-response curve of a new drug with respect to toxic side effects and, in particular, to estimate the maximum tolerated dose (MTD). In this paper we take a Bayesian approach to the problem of making inferences about the MTD. Working with broad classes of priors, we obtain the posterior distribution of the MTD and study its properties. We also address the question of providing updated assessments of the risk of toxicity for new patients entering the study at a specific dose level. These assessments would be useful in deciding issues of study management and ethics. Our analysis pays particular attention to the sensitivity of the inferences and risk assessments to the choice of prior and the choice of model for the dose-response relationship.

Algorithms↗

A Bayesian approach to the estimation of ancestral genome arrangements.

We describe a Bayesian approach to estimate phylogeny and ancestral genome arrangements on the basis of genome arrangement data using a model in which gene inversion is the sole mechanism of change. While we have described a similar method to estimate phylogenetic relationships in the statistics literature, the novel contribution of the present work is the description of a method to compute probability distributions of ancestral genome arrangements. We assess the robustness of posterior distributions to different specifications of prior distributions and provide an empirical means to selecting a prior distribution. We note that parsimony approaches to ancestral reconstruction in the literature focus on the development of computationally efficient algorithms for searching for optimal ancestral genome arrangements, but, unlike Bayesian approaches, do not include assessment of uncertainty in these estimates. We compare and contrast a Bayesian approach with a parsimony approach to infer phylogenies and ancestral arrangements from genome arrangement data by re-analyzing a number of previously published data sets.

Algorithms↗

Evolution of carnivory in Lentibulariaceae and the Lamiales.

As a basis for analysing the evolution of the carnivorous syndrome in Lentibulariaceae (Lamiales), phylogenetic reconstructions were conducted based on coding and non-coding chloroplast DNA (matK gene and flanking trnK intron sequences, totalling about 2.4 kb). A dense taxon sampling including all other major lineages of Lamiales was needed since the closest relatives of Lentibulariaceae and the position of "proto-carnivores" were unknown. Tree inference using maximum parsimony, maximum likelihood, and Bayesian approaches resulted in fully congruent topologies within Lentibulariaceae, whereas relationships among the different lineages of Lamiales were only congruent between likelihood and Bayesian optimizations. Lentibulariaceae and their three genera (Pinguicula, Genlisea, and Utricularia) are monophyletic, with Pinguicula being sister to a Genlisea-Utricularia clade. Likelihood and Bayesian trees converge on Bignoniaceae as sister to Lentibulariaceae, albeit lacking good support. The "proto-carnivores" (Byblidaceae, Martyniaceae) are found in different positions among other Lamiales but not as sister to the carnivorous Lentibulariaceae, which is also supported by Khishino-Hasegawa tests. This implies that carnivory and its preliminary stages ("proto-carnivores") independently evolved more than once among Lamiales. Ancestral states of structural characters connected to the carnivorous syndrome are reconstructed using the molecular tree, and a hypothesis on the evolutionary pathway of the carnivorous syndrome in Lentibulariaceae is presented. Extreme DNA mutational rates found in Utricularia and Genlisea are shown to correspond to their unusual nutritional specialization, thereby hinting at a marked degree of carnivory in these two genera.

Animals↗

Bayesian error analysis model for reconstructing transcriptional regulatory networks.

Transcription regulation is a fundamental biological process, and extensive efforts have been made to dissect its mechanisms through direct biological experiments and regulation modeling based on physical-chemical principles and mathematical formulations. Despite these efforts, transcription regulation is yet not well understood because of its complexity and limitations in biological experiments. Recent advances in high throughput technologies have provided substantial amounts and diverse types of genomic data that reveal valuable information on transcription regulation, including DNA sequence data, protein-DNA binding data, microarray gene expression data, and others. In this article, we propose a Bayesian error analysis model to integrate protein-DNA binding data and gene expression data to reconstruct transcriptional regulatory networks. There are two unique aspects to this proposed model. First, transcription is modeled as a set of biochemical reactions, and a linear system model with clear biological interpretation is developed. Second, measurement errors in both protein-DNA binding data and gene expression data are explicitly considered in a Bayesian hierarchical model framework. Model parameters are inferred through Markov chain Monte Carlo. The usefulness of this approach is demonstrated through its application to infer transcriptional regulatory networks in the yeast cell cycle.

Algorithms↗

Genetic consequences of sequential founder events by an island-colonizing bird.

The importance of founder events in promoting evolutionary changes on islands has been a subject of long-running controversy. Resolution of this debate has been hindered by a lack of empirical evidence from naturally founded island populations. Here we undertake a genetic analysis of a series of historically documented, natural colonization events by the silvereye species-complex (Zosterops lateralis), a group used to illustrate the process of island colonization in the original founder effect model. Our results indicate that single founder events do not affect levels of heterozygosity or allelic diversity, nor do they result in immediate genetic differentiation between populations. Instead, four to five successive founder events are required before indices of diversity and divergence approach that seen in evolutionarily old forms. A Bayesian analysis based on computer simulation allows inferences to be made on the number of effective founders and indicates that founder effects are weak because island populations are established from relatively large flocks. Indeed, statistical support for a founder event model was not significantly higher than for a gradual-drift model for all recently colonized islands. Taken together, these results suggest that single colonization events in this species complex are rarely accompanied by severe founder effects, and multiple founder events and/or long-term genetic drift have been of greater consequence for neutral genetic diversity.

Animals↗

Design and analysis of admixture mapping studies.

Admixture between populations originating on different continents can be exploited to detect disease susceptibility loci at which risk alleles are distributed differentially between these populations. We first examine the statistical power and mapping resolution of this approach in the limiting situation in which gamete admixture and locus ancestry are measured without uncertainty. We show that, for a rare disease, the most efficient design is to study affected individuals only. In a typical African American population (two-way admixture proportions 0.8/0.2, ancestry crossover rate 2 per 100 cM), a study of 800 affected individuals has 90% power to detect at P values <10(-5) a locus that generates a risk ratio of 2 between populations, with an expected mapping resolution (size of 95% confidence region for the position of the locus) of 4 cM. In practice, to infer locus ancestry from marker data requires Bayesian computationally intensive methods, as implemented in the program ADMIXMAP. Affected-only study designs require strong prior information on the frequencies of each allele given locus ancestry. We show how data from unadmixed and admixed populations can be combined to estimate these ancestry-specific allele frequencies within the admixed population under study, allowing for variation between allele frequencies in unadmixed and admixed populations. Using simulated data based on the genetic structure of the African American population, we show that 60% of information can be extracted in a test for linkage using markers with an ancestry information content of 36% at 3-cM spacing. As in classic linkage studies, the most efficient strategy is to use markers at a moderate density for an initial genome search and then to saturate regions of putative linkage with additional markers, to extract nearly all information about locus ancestry.

Black People↗

Identifying multigenic modules under selection in the tumor genome.

MOTIVATION: Genomic alterations in cancer arise from selective pressures acting on hallmark molecular modules, layered over a background of random mutagenic events. Methods to detect selection at the level of modules, as opposed to genes or nucleotides, are relatively underdeveloped. RESULTS: Here we present CanSRMaPP (Cancer Selection Recovery by Maximum Posterior Probability), a Bayesian model of the cancer genome that infers mutational selection on single genes and multi-genic modules while simultaneously modeling background events. Applying CanSRMaPP to lung adenocarcinoma genomes, we identify positive selection on 63 modules, yielding a model that parsimoniously explains the observed pattern of genetic alterations observed in new cancer cohorts. We further show that CanSRMaPP is adaptable to more tumor types and to alternative module definitions. We show that these modules serve as an effective scaffold for translating the cancer genome to molecular states, with prediction of cancer biomarker status as demonstration. AVAILABILITY: CanSRMaPP is freely available on GitHub. SUPPLEMENTARY INFORMATION: Supplementary Figs. S1-5, Supplementary Tables S1-5, and Supplementary Notes 1 and 2 are available at Bioinformatics online.

Journal Article↗