Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Untangling the Arisaema enigma: Investigating the complex evolutionary history and species relationships in North American Arisaema.

PREMISE: The evolutionary history of morphologically variable plant groups is often obscured by cryptic diversity, morphological convergence, and limited genetic data. Arisaema, a diverse genus within Araceae, exemplifies these challenges. Although some taxonomic treatments recognize only two species of North American Arisaema (A. dracontium and A. triphyllum), other studies have identified morphologically distinct groups within both taxa. Here, we reconstructed evolutionary relationships in North American Arisaema, assessed genetic structure and admixture, and tested the monophyly of proposed species. METHODS: We used 2b-RAD sequencing to generate genome-wide SNP data for 146 samples from 31 populations across the eastern United States. Phylogenetic relationships were inferred using maximum-likelihood and Bayesian approaches. Population structure and admixture were assessed using the program structure and principal component analysis (PCA). RESULTS: Both the Arisaema triphyllum and A. dracontium complexes formed well-supported monophyletic groups. Within the A. dracontium complex, we recovered three monophyletic lineages: A. dracontium, A. calciphilum, and A. macrospathum. In the A. triphyllum complex, A. quinatum, A. stewardsonii, and A. allegheniense consistently formed distinct groups. Relationships between A. pusillum and A. acuminatum, and among A. triphyllum s.s., A. purpurascens, and A. striatum were less clearly resolved, likely due to recent or incomplete divergence, gene flow, or polyploidy. CONCLUSIONS: The results support the monophyly of multiple newly proposed taxa within North American Arisaema, but additional sampling across the species' ranges is needed to fully resolve species boundaries. Our study provides the first evolutionary framework for this group, providing a foundation for future ecological, taxonomic, and conservation research in the genus.

Araceae↗

A Bayesian fixed effects analysis of the Mantel-Haenszel model applied to meta-analysis.

When performing a meta analysis, it is often necessary to combine results from several 2 x 2 contingency tables. The Mantel-Haenszel model assumes a common measure of association between the treatment and outcome variables across the tables. A Bayesian method is described for drawing inferences regarding the measure of association, for checking the plausibility of the Mantel-Haenszel model, and for drawing inferences regarding the success rates for the individual studies. While the methodology is readily extendable to random effects models, a fixed effects approach avoids the complex statistical modelling of a mixture distribution which is required for the good application of random effects models.

Anti-Bacterial Agents↗

Bayesian analysis of population structure based on linked molecular information.

The Bayesian model-based approach to inferring hidden genetic population structures using multilocus molecular markers has become a popular tool within certain branches of biology. In particular, it has been shown that heterogeneous data arising from genetically dissimilar latent groups of individuals can be effectively modelled using an unsupervised classification formulation. However, most currently employed models ignore potential linkage within the employed molecular information, and can therefore lead to biased inferences under certain circumstances. Utilizing the general theory of graphical models, we develop a framework that accounts for dependences both within linked molecular marker loci and DNA sequence data. Due to a high level of sequence conservation among eukaryotic species, the latter aspect is particularly relevant for analyzing rapidly evolving microbial species. The advantages of incorporating the dependence due to linkage in the classification models are illustrated by analyses of both simulated data and real samples of Bacillus cereus.

Bacillus cereus↗

The hepatitis C virus epidemic in Cameroon: genetic evidence for rapid transmission between 1920 and 1960.

Hepatitis C virus (HCV) infection in Cameroon is characterized by widespread seropositivity and great virus genetic diversity (3 genotypes and over 10 subtypes). A total of 244 HCV NS5B sequences of 382-405 bp long (95 type 1, 58 type 2, and 91 type 4) were phylogenetically analyzed to estimate the history of the HCV epidemic in Cameroon. The newly developed Bayesian coalescent approach was used to infer the history of each HCV type. The estimated dates of the most recent common ancestors (MRCA) for genotypes 1 (1500; 95% confidence interval (95% CI): 1300-1650) and 4 (1500; 95% CI: 1350-1700) were in the same range, while the date for genotype 2 MRCA (1600; 95% CI: 1400-1750) was slightly more recent. The mean genetic distance between HCV genotype 1 sequences was greater than that of HCV type 4 sequences, itself greater than that of HCV type 2 sequences. The initial infected populations of all three genotypes did not grow until recently, when they grew exponentially. The growth rate has now begun to slow, with a less steep exponential growth curve. The period of exponential growth of all the three genotypes was between 1920 and 1960. These results (i) confirm that HCV genotypes 1 and 4 have produced long-term endemics, (ii) suggest that genotype 2 was introduced into Cameroon more recently, and (iii) indicate that the exponential spread of the three genotypes between 1920 and 1960 coincided with the mass campaign against trypanosomiasis and mass vaccinations in Cameroon.

Bayes Theorem↗

A mitochondrial phylogeny of the rainforest skink genus Saproscincus, Wells and Wellington (1984).

The phylogenetic relationships and historical biogeography of 10 currently described rainforest skinks in the genus Saproscincus were investigated using mitochondrial protein-coding ND4 and ribosomal RNA 16S genes. A robust phylogeny is inferred using both maximum likelihood and Bayesian analysis, with all inter-specific nodes strongly supported when datasets are combined. The phylogeny supports the recognition of two major lineages (northern and southern), each of which comprises two divergent clades. Both northern and southern lineages have comparably divergent representatives in mid-east Queensland (MEQ), providing further molecular evidence for the importance of two major biogeographic breaks, the St. Lawrence gap and Burdekin gap separating MEQ from southern and northern counterparts respectively. Vicariance associated with the fragmentation and contraction of temperate rainforest during the mid-late Miocene epoch underpins the deep divergence between morphologically conservative lineages in at least three instances. In contrast, one species, Saproscincus oriarus, shows very low sequence divergence but distinct morphological and ecological differentiation from its allopatric sister clade within Saproscincus mustelinus. These results suggest that while vicariance has played a prominent role in diversification and historical biogeography of Saproscincus, divergent selection may also be important.

Animals↗

Singularities in mixture models and upper bounds of stochastic complexity.

A learning machine which is a mixture of several distributions, for example, a gaussian mixture or a mixture of experts, has a wide range of applications. However, such a machine is a non-identifiable statistical model with a lot of singularities in the parameter space, hence its generalization property is left unknown. Recently an algebraic geometrical method has been developed which enables us to treat such learning machines mathematically. Based on this method, this paper rigorously proves that a mixture learning machine has the smaller Bayesian stochastic complexity than regular statistical models. Since the generalization error of a learning machine is equal to the increase of the stochastic complexity, the result of this paper shows that the mixture model can attain the more precise prediction than regular statistical models if Bayesian estimation is applied in statistical inference.

Bayes Theorem↗

Bayesian estimation of the number of inversions in the history of two chromosomes.

We present a Bayesian approach to the problem of inferring the history of inversions separating homologous chromosomes from two different species. The method is based on Markov Chain Monte Carlo (MCMC) and takes full advantage of all the information from marker order. We apply the method both to simulated data and to two real data sets. For the simulated data, we show that the MCMC method provides accurate estimates of the true posterior distributions and in the analysis of the real data we show that the most likely number of inversions in some cases is considerably larger than estimates obtained based on the parsimony inferred number of inversions. Indeed, in the case of the Drosophila repleta-D. melanogaster comparison, the lower boundary of a 95% highest posterior density credible interval for the number of inversions is considerably larger than the most parsimonious number of inversions.

Animals↗

Multivariate refutation of aetiological hypotheses in non-experimental epidemiology.

Extension of Karl Popper's logic of refutation from the realm of contingency tables to multivariate modelling leads to the conclusion that rigorously scientific multivariate analysis in non-experimental epidemiology differs from the traditional quasi-scientific approach. Instead of aiming for high sensitivity in detecting aetiological agents, the goal in refutation is high specificity--to give the best defence of the 'innocence' of every exposure hypothesized as being a cause. Instead of 'forward selection' or 'backward elimination', multivariate refutation uses the method of 'forward elimination'. This entails a likelihood approach (which may be complemented by, but should be demarcated from, Bayesian methods) not only for statistical inference but also, by analogy, for study design and conduct: one starts with the conclusion (the estimate or hypothesis) and works backwards to the observations (the likelihood of the data or the design of the study). Differences in practice can sometimes be large, as illustrated by a study of hypothesized triggers of myocardial infarction. Multivariate refutation should replace the concept of multivariate modelling in non-experimental epidemiology.

Bayes Theorem↗

A molecular time-scale for eukaryote evolution recalibrated with the continuous microfossil record.

Recent attempts to establish a molecular time-scale of eukaryote evolution failed to provide a congruent view on the timing of the origin and early diversification of eukaryotes. The major discrepancies in molecular time estimates are related to questions concerning the calibration of the tree. To limit these uncertainties, we used here as a source of calibration points the rich and continuous microfossil record of dinoflagellates, diatoms and coccolithophorids. We calibrated a small-subunit ribosomal RNA tree of eukaryotes with four maximum and 22 minimum time constraints. Using these multiple calibration points in a Bayesian relaxed molecular clock framework, we inferred that the early radiation of eukaryotes occurred near the Mesoproterozoic-Neoproterozoic boundary, about 1100 million years ago. Our results indicate that most Proterozoic fossils of possible eukaryotic origin cannot be confidently assigned to extant lineages and should therefore not be used as calibration points in molecular dating.

Amoeba↗

Hunting drug targets by systems-level modeling of gene expression profiles.

Structural learning of Bayesian networks applied to sets of genome-wide expression patterns has been recently discovered as a potentially useful tool for the systems-level statistical description of gene interactions. We train and analyze Bayesian networks with the goal of inferring biological aspects of gene function. Our two-component approach focuses on supporting the drug discovery process by identifying genes with central roles for the network operation, which could act as drug targets. The first component, referred to as scale-free analysis, uses topological measures of the network-related to a high-traffic load of genes-as estimators for their functional importance. The second component, referred to as generative inverse modeling, is a method of estimating the effect of a simulated drug treatment or mutation on the global state of the network, as measured in the expression profile. We show for a dataset from acute lymphoblastic leukemia patients that both approaches are suitable for finding genes with central cellular functions. In addition, generative inverse modeling correctly identifies a known oncogene in a purely data-driven way.

Algorithms↗

A dynamic causal modeling study on category effects: bottom-up or top-down mediation?

In this study, we combined functional magnetic resonance imaging (fMRI) and dynamic causal modeling (DCM) to investigate whether object category effects in the occipital anti temporal cortex are mediated by inputs from early visual cortex or parietal regions. Resolving this issue may provide anatomical constraints on theories of category specificity--which make different assumptions about the underlying neurophysiology. The data were acquired by Ishai, Ungerleider, Martin, Schouten, and Haxby (1999, 2000) and provided by the National fiNRI Data Center (http://www.fmridcc.org). The original authors used a conventional analysis to estimate differential effects in the occipital anti temporal cortex in response to pictures of chairs, faces, and houses. We extended this approach by estimating neuronal interactions that mediate category effects using DCM. DCM uses a Bayesian framework to estimate and make inferences about the influence that one region exerts over another and how this is affected by experimental changes. DCM differs from previous approaches to brain connectivity, such as multivariate autoregressive models and structural equation modeling, as it assumes that the observed hemodynamic responses are driven by experimental changes rather than endogenous noise. DCM therefore brings the analysis of brain connectivity much closer to the analysis of regionally specific effects usually applied to functional imaging data. We used DCM to estimate the influence that V3 and the superior/inferior parietal cortex exerted over category-responsive regions and how this was affected by the presentation of houses, faces, and chairs. We found that category effects in occipital and temporal cortex were mediated by inputs from early visual cortex. In contrast,the connectivity from the superior/inferior parietal area to the category-responsive areas was unaffected by the presentation of chairs, faces, or houses. These findings indicate that category effects in the occipital and temporal cortex can be mediated by bottom-up mechanisms-a finding that needs to be embraced by models of category specificity.

Brain↗

Bayesian estimation of genomic distance.

We present a Bayesian approach to the problem of inferring the number of inversions and translocations separating two species. The main reason for developing this method is that it will allow us to test hypotheses about the underlying mechanisms, such as the distribution of inversion track lengths or rate constancy among lineages. Here, we apply these methods to comparative maps of eggplant and tomato, human and cat, and human and cattle with 170, 269, and 422 markers, respectively. In the first case the most likely number of events is larger than the parsimony value. In the last two cases the parsimony solutions have very small probability.

Animals↗

Model weights and the foundations of multimodel inference.

Statistical thinking in wildlife biology and ecology has been profoundly influenced by the introduction of AIC (Akaike's information criterion) as a tool for model selection and as a basis for model averaging. In this paper, we advocate the Bayesian paradigm as a broader framework for multimodel inference, one in which model averaging and model selection are naturally linked, and in which the performance of AIC-based tools is naturally evaluated. Prior model weights implicitly associated with the use of AIC are seen to highly favor complex models: in some cases, all but the most highly parameterized models in the model set are virtually ignored a priori. We suggest the usefulness of the weighted BIC (Bayesian information criterion) as a computationally simple alternative to AIC, based on explicit selection of prior model probabilities rather than acceptance of default priors associated with AIC. We note, however, that both procedures are only approximate to the use of exact Bayes factors. We discuss and illustrate technical difficulties associated with Bayes factors, and suggest approaches to avoiding these difficulties in the context of model selection for a logistic regression. Our example highlights the predisposition of AIC weighting to favor complex models and suggests a need for caution in using the BIC for computing approximate posterior model weights.

Animals↗

Oh brother, where art thou? A Bayes factor test for recombination with uncertain heritage.

Current methods to identify recombination between subtypes of human immunodeficiency virus 1 (HIV-1) fall into a sequential testing trap, in which significance is assessed conditional on parental representative sequences and crossover points (COPs) that maximize the same test statistic. We overcame this shortfall by testing for recombination while inferring parental heritage and COPs using an extended Bayesian multiple change-point model. The model assumes that aligned molecular sequence data consist of an unknown number of contiguous segments that may support alternative topologies or varying evolutionary pressures. We allowed for heterogeneity in the substitution process and specifically tested for intersubtype recombination using Bayes factors. We also developed a new class of priors to assess significance across a wide range of support for recombination in the data. We applied our method to three putative gag gene recombinants. HIV-1 isolate RW024 decisively supported recombination with an inferred parental heritage of AD and a COP 95% Bayesian credible interval of (1,152, 1,178) using the HXB2 numbering scheme. HIV-1 isolate VI557 barely supported recombination. HIV-1 isolate RF decisively rejected recombination as expected, given that the sequence is commonly used as a reference sequence for subtype B. We employed scaled regeneration quantile plots to assess convergence and found this approach convenient to use even for our variable dimensional model parameter space.

Algorithms↗

Phylogeny and historical biogeography of African ground squirrels: the role of climate change in the evolution of Xerus.

We used phylogenetic and phylogeographical methods to infer relationships among African ground squirrels of the genus Xerus. Using Bayesian, maximum-parsimony, nested clade and coalescent analyses of cytochrome b sequences, we inferred interspecific relationships, evaluated the specific distinctness of Cape (Xerus inauris) and mountain (Xerus princeps) ground squirrels, and tested hypotheses for historical patterns of gene flow within X. inauris. The inferred phylogeny supports the hypothesized existence of an 'arid corridor' from the Horn of Africa to the Cape region. Although doubts have been raised regarding the specific distinctness of X. inauris and X. princeps, our analyses show that each represents a distinct well-supported, monophyletic lineage. Xerus inauris includes three major clades, two of which are geographically restricted. The distributions of X. inauris populations are concordant with divergences within and disjunctions between other taxa, which have been interpreted as results of Plio-Pleistocene climate cycles. Nested clade analysis, coalescent analyses, and analyses of genetic structure support allopatric fragmentation as the cause of the deep divergences within this species.

Africa↗

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem↗

Applying dynamic Bayesian networks to perturbed gene expression data.

BACKGROUND: A central goal of molecular biology is to understand the regulatory mechanisms of gene transcription and protein synthesis. Because of their solid basis in statistics, allowing to deal with the stochastic aspects of gene expressions and noisy measurements in a natural way, Bayesian networks appear attractive in the field of inferring gene interactions structure from microarray experiments data. However, the basic formalism has some disadvantages, e.g. it is sometimes hard to distinguish between the origin and the target of an interaction. Two kinds of microarray experiments yield data particularly rich in information regarding the direction of interactions: time series and perturbation experiments. In order to correctly handle them, the basic formalism must be modified. For example, dynamic Bayesian networks (DBN) apply to time series microarray data. To our knowledge the DBN technique has not been applied in the context of perturbation experiments. RESULTS: We extend the framework of dynamic Bayesian networks in order to incorporate perturbations. Moreover, an exact algorithm for inferring an optimal network is proposed and a discretization method specialized for time series data from perturbation experiments is introduced. We apply our procedure to realistic simulations data. The results are compared with those obtained by standard DBN learning techniques. Moreover, the advantages of using exact learning algorithm instead of heuristic methods are analyzed. CONCLUSION: We show that the quality of inferred networks dramatically improves when using data from perturbation experiments. We also conclude that the exact algorithm should be used when it is possible, i.e. when considered set of genes is small enough.

Algorithms↗

Gene network inference from incomplete expression data: transcriptional control of hematopoietic commitment.

MOTIVATION: The topology and function of gene regulation networks are commonly inferred from time series of gene expression levels in cell populations. This strategy is usually invalid if the gene expression in different cells of the population is not synchronous. A promising, though technically more demanding alternative is therefore to measure the gene expression levels in single cells individually. The inference of a gene regulation network requires knowledge of the gene expression levels at successive time points, at least before and after a network transition. However, owing to experimental limitations a complete determination of the precursor state is not possible. RESULTS: We investigate a strategy for the inference of gene regulatory networks from incomplete expression data based on dynamic Bayesian networks. This permits prediction of the number of experiments necessary for network inference depending on parameters including noise in the data, prior knowledge and limited attainability of initial states. Our strategy combines a gradual 'Partial Learning' approach based solely on true experimental observations for the network topology with expectation maximization for the network parameters. We illustrate our strategy by extensive computer simulations in a high-dimensional parameter space in a simulated single-cell-based example of hematopoietic stem cell commitment and in random networks of different sizes. We find that the feasibility of network inferences increases significantly with the experimental ability to force the system into different initial network states, with prior knowledge and with noise reduction. AVAILABILITY: Source code is available under: www.izbi.uni-leipzig.de/services/NetwPartLearn.html SUPPLEMENTARY INFORMATION: Supplementary Data are available at Bioinformatics online.

Algorithms↗