Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Phylogenetic inference in Rafflesiales: the influence of rate heterogeneity and horizontal gene transfer.

BACKGROUND: The phylogenetic relationships among the holoparasites of Rafflesiales have remained enigmatic for over a century. Recent molecular phylogenetic studies using the mitochondrial matR gene placed Rafflesia, Rhizanthes and Sapria (Rafflesiaceae s. str.) in the angiosperm order Malpighiales and Mitrastema (Mitrastemonaceae) in Ericales. These phylogenetic studies did not, however, sample two additional groups traditionally classified within Rafflesiales (Apodantheaceae and Cytinaceae). Here we provide molecular phylogenetic evidence using DNA sequence data from mitochondrial and nuclear genes for representatives of all genera in Rafflesiales. RESULTS: Our analyses indicate that the phylogenetic affinities of the large-flowered clade and Mitrastema, ascertained using mitochondrial matR, are congruent with results from nuclear SSU rDNA when these data are analyzed using maximum likelihood and Bayesian methods. The relationship of Cytinaceae to Malvales was recovered in all analyses. Relationships between Apodanthaceae and photosynthetic angiosperms varied depending upon the data partition: Malvales (3-gene), Cucurbitales (matR) or Fabales (atp1). The latter incongruencies suggest that horizontal gene transfer (HGT) may be affecting the mitochondrial gene topologies. The lack of association between Mitrastema and Ericales using atp1 is suggestive of HGT, but greater sampling within eudicots is needed to test this hypothesis further. CONCLUSIONS: Rafflesiales are not monophyletic but composed of three or four independent lineages (families): Rafflesiaceae, Mitrastemonaceae, Apodanthaceae and Cytinaceae. Long-branch attraction appears to be misleading parsimony analyses of nuclear small-subunit rDNA data, but model-based methods (maximum likelihood and Bayesian analyses) recover a topology that is congruent with the mitochondrial matR gene tree, thus providing compelling evidence for organismal relationships. Horizontal gene transfer appears to be influencing only some taxa and some mitochondrial genes, thus indicating that the process is acting at the single gene (not whole genome) level.

Bayes Theorem↗

Gene networks as a tool to understand transcriptional regulation.

Gene regulatory networks, or simply gene networks (GNs), have shown to be a promising approach that the bioinformatics community has been developing for studying regulatory mechanisms in biological systems. GNs are built from the genome-wide high-throughput gene expression data that are often available from DNA microarray experiments. Conceptually, GNs are (un)directed graphs, where the nodes correspond to the genes and a link between a pair of genes denotes a regulatory interaction that occurs at transcriptional level. In the present study, we had two objectives: 1) to develop a framework for GN reconstruction based on a Bayesian network model that captures direct interactions between genes through nonparametric regression with B-splines, and 2) to demonstrate the potential of GNs in the analysis of expression data of a real biological system, the yeast pheromone response pathway. Our framework also included a number of search schemes to learn the network. We present an intuitive notion of GN theory as well as the detailed mathematical foundations of the model. A comprehensive analysis of the consistency of the model when tested with biological data was done through the analysis of the GNs inferred for the yeast pheromone pathway. Our results agree fairly well with what was expected based on the literature, and we developed some hypotheses about this system. Using this analysis, we intended to provide a guide on how GNs can be effectively used to study transcriptional regulation. We also discussed the limitations of GNs and the future direction of network analysis for genomic data. The software is available upon request.

Bayes Theorem↗

Bayesian segmental models with multiple sequence alignment profiles for protein secondary structure and contact map prediction.

In this paper, we develop a segmental semi-Markov model (SSMM) for protein secondary structure prediction which incorporates multiple sequence alignment profiles with the purpose of improving the predictive performance. The segmental model is a generalization of the hidden Markov model where a hidden state generates segments of various length and secondary structure type. A novel parameterized model is proposed for the likelihood function that explicitly represents multiple sequence alignment profiles to capture the segmental conformation. Numerical results on benchmark data sets show that incorporating the profiles results in substantial improvements and the generalization performance is promising. By incorporating the information from long range interactions in beta-sheets, this model is also capable of carrying out inference on contact maps. This is an important advantage of probabilistic generative models over the traditional discriminative approach to protein secondary structure prediction. The Web server of our algorithm and supplementary materials are available at http://public.kgi.edu/-wild/bsm.html.

Algorithms↗

Bayesian meta-analysis for longitudinal data models using multivariate mixture priors.

We propose a class of longitudinal data models with random effects that generalizes currently used models in two important ways. First, the random-effects model is a flexible mixture of multivariate normals, accommodating population heterogeneity, outliers, and nonlinearity in the regression on subject-specific covariates. Second, the model includes a hierarchical extension to allow for meta-analysis over related studies. The random-effects distributions are decomposed into one part that is common across all related studies (common measure), and one part that is specific to each study and that captures the variability intrinsic between patients within the same study. Both the common measure and the study-specific measures are parameterized as mixture-of-normals models. We carry out inference using reversible jump posterior simulation to allow a random number of terms in the mixtures. The sampler takes advantage of the small number of entertained models. The motivating application is the analysis of two studies carried out by the Cancer and Leukemia Group B (CALGB). In both studies, we record for each patient white blood cell counts (WBC) over time to characterize the toxic effects of treatment. The WBCs are modeled through a nonlinear hierarchical model that gathers the information from both studies.

Bayes Theorem↗

Adaptive diagnosis in distributed systems.

Real-time problem diagnosis in large distributed computer systems and networks is a challenging task that requires fast and accurate inferences from potentially huge data volumes. In this paper, we propose a cost-efficient, adaptive diagnostic technique called active probing. Probes are end-to-end test transactions that collect information about the performance of a distributed system. Active probing uses probabilistic reasoning techniques combined with information-theoretic approach, and allows a fast online inference about the current system state via active selection of only a small number of most-informative tests. We demonstrate empirically that the active probing scheme greatly reduces both the number of probes (from 60% to 75% in most of our real-life applications), and the time needed for localizing the problem when compared with nonadaptive (preplanned) probing schemes. We also provide some theoretical results on the complexity of probe selection, and the effect of "noisy" probes on the accuracy of diagnosis. Finally, we discuss how to model the system's dynamics using dynamic Bayesian networks (DBNs), and an efficient approximate approach called sequential multifault; empirical results demonstrate clear advantage of such approaches over "static" techniques that do not handle system's changes.

Algorithms↗

An innovative application of Bayesian disease mapping methods to patient safety research: a Canadian adverse medical event study.

Recently developed disease mapping and ecological regression methods have become important techniques in studies of disease epidemiology and in health services research. This increase in importance is partially a result of the development of Bayesian statistical methodologies that make it possible to study associations between health problems and risk factors at an aggregate (i.e. areal) level while taking into account such matters as unmeasured confounding and spatial relationships. In this paper we present a demonstration of the joint use of empirical Bayes (EB) and full Bayesian inferential techniques in a small area study of adverse medical events (also known as 'iatrogenic injury') in British Columbia, Canada. In particular, we illustrate a unified Bayesian hierarchical spatial modelling framework that enables simultaneous examinations of potential associations between adverse medical event occurrence and regional characteristics, age effects, residual variation and spatial autocorrelation. We propose an analytic strategy for complementary use of EB and FB inferential techniques for risk assessment and model selection, presenting an EB-FB combined approach that draws on the strengths of each method while minimizing inherent weaknesses. The work was motivated by the need to explore relatively efficient ways to analyse regional variations of health services outcomes and resource utilization when a considerable amount of statistical modelling and inference are required.

Adolescent↗

Combining phylogenetic motif discovery and motif clustering to predict co-regulated genes.

MOTIVATION: We present a sequence-based framework and algorithm PHYLOCLUS for predicting co-regulated genes. In our approach, de novo discovery methods are used to find motifs conserved by evolution and then a Bayesian hierarchical clustering model is used to cluster these motifs, thereby grouping together genes that are putatively co-regulated. Our clustering procedure allows both the number of clusters and the motif width within each cluster to be unknown. RESULTS: We use our framework to predict co-regulated genes in the bacterium Bacillus subtilis using six other closely related bacterial species. Our predicted motifs and gene clusters are validated using several external sources and significant clusters are examined in detail. An extension to the discovery and clustering of two-block motifs can be used for inference about synergistic binding relationships between transcription factors. AVAILABILITY: Software and Supplementary Materials can be downloaded at http://stat.wharton.upenn.edu/~stjensen/research/phyloclus.html or http://www.fas.harvard.edu/~junliu/phyloclus.html CONTACT: stjensen@wharton.upenn.edu.

Algorithms↗

An elusive paleodemography? A comparison of two methods for estimating the adult age distribution of deaths at late Classic Copan, Honduras.

Comparison of different adult age estimation methods on the same skeletal sample with unknown ages could forward paleodemographic inference, while researchers sort out various controversies. The original aging method for the auricular surface (Lovejoy et al., 1985a) assigned an age estimation based on several separate characteristics. Researchers have found this original method hard to apply. It is usually forgotten that before assigning an age, there was a seriation, an ordering of all available individuals from youngest to oldest. Thus, age estimation reflected the place of an individual within its sample. A recent article (Buckberry and Chamberlain, 2002) proposed a revised method that scores theses various characteristics into age stages, which can then be used with a Bayesian method to estimate an adult age distribution for the sample. Both methods were applied to the adult auricular surfaces of a Pre-Columbian Maya skeletal population from Copan, Honduras and resulted in age distributions with significant numbers of older adults. However, contrary to the usual paleodemographic distribution, one Bayesian estimation based on uniform prior probabilities yielded a population with 57% of the ages at death over 65, while another based on a high mortality life table still had 12% of the individuals aged over 75 years. The seriation method yielded an age distribution more similar to that known from preindustrial historical situations, without excessive longevity of adults. Paleodemography must still wrestle with its elusive goal of accurate adult age estimation from skeletons, a necessary base for demographic study of past populations.

Age Determination by Skeleton↗

Decoding non-unique oligonucleotide hybridization experiments of targets related by a phylogenetic tree.

MOTIVATION: The reliable identification of presence or absence of biological agents ("targets"), such as viruses or bacteria, is crucial for many applications from health care to biodiversity. If genomic sequences of targets are known, hybridization reactions between oligonucleotide probes and targets performed on suitable DNA microarrays will allow to infer presence or absence from the observed pattern of hybridization. Targets, for example all known strains of HIV, are often closely related and finding unique probes becomes impossible. The use of non-unique oligonucleotides with more advanced decoding techniques from statistical group testing allows to detect known targets with great success. Of great relevance, however, is the problem of identifying the presence of previously unknown targets or of targets that evolve rapidly. RESULTS: We present the first approach to decode hybridization experiments using non-unique probes when targets are related by a phylogenetic tree. Using a Bayesian framework and a Markov chain Monte Carlo approach we are able to identify over 94% of known targets and assign up to 70% of unknown targets to their correct clade in hybridization simulations on biological and simulated data. AVAILABILITY: Software implementing the method described in this paper and datasets are available from http://algorithmics.molgen.mpg.de/probetrees.

Algorithms↗

Bayesian statistics as applied to hypertension diagnosis.

This paper deals with the following hypertension diagnoses: essential hypertension and five types of secondary hypertension: fibrodysplasic renal artery stenosis, atheromatous renal artery stenosis, Conn's syndrome, renal cystic disease, and pheochromocytoma. Only blood pressures, general information and general biochemical data are taken into account. Nineteen items were finally selected, by statistical investigation of experimental data, as being both discriminative and independent. The marginal density distributions of every item, and then joint density distribution functions were determined within six types of hypertension. The frequency of a given hypertension type within the hypertensive patients was used as prior probability of this state. The loss matrix was established by medical arguments. The expected loss corresponding to six possible decisions could thus be calculated for all cases. Both the ratio of secondary hypertensions that could be inferred from our set of data (not including the results of complementary tests) and that of correct "essential" hypertension diagnosis proved to be satisfactory.

Bayes Theorem↗

Early postmarketing drug safety surveillance: data mining points to consider.

BACKGROUND: Computer-assisted data mining algorithms (DMAs) are being studied to screen spontaneous reporting databases for signals of novel adverse events. The performance characteristics and optimum deployment of these techniques remain to be established. OBJECTIVE: To explore issues in the practical evaluation and deployment of DMAs by comparing findings from an empirical Bayesian DMA with those from a traditional drug safety surveillance program. METHODS: Published findings from early postmarketing safety surveillance of thalidomide were compared with findings from an empirical Bayesian DMA. Differential results were used to explore practical issues in the evaluation and deployment of DMAs. RESULTS: Most adverse events highlighted by each method were compatible with the product labeling or natural history/complications of reported treatment indications. Traditional surveillance highlighted 4 potentially serious and unexpected adverse events (Stevens-Johnson syndrome, toxic epidermal necrolysis, seizures, skin ulcers) warranting labeling amendments or close monitoring. None of these adverse event terms generated a signal using the DMA. CONCLUSIONS: The DMA would not have enhanced early postmarketing surveillance in this particular setting. While the results cannot be used to draw inferences about the global performance of DMAs, they illustrate the following: (1) DMA performance may be highly situation dependent; (2) over-reliance on these methods may have deleterious consequences, especially with so-called "designated medical events"; and (3) the most appropriate selection of pharmacovigilance tools needs to be tailored to each situation, being mindful of the numerous factors that may influence comparative performance and incremental utility of DMAs.

Algorithms↗

Control of confounding of genetic associations in stratified populations.

To control for hidden population stratification in genetic-association studies, statistical methods that use marker genotype data to infer population structure have been proposed as a possible alternative to family-based designs. In principle, it is possible to infer population structure from associations between marker loci and from associations of markers with the trait, even when no information about the demographic background of the population is available. In a model in which the total population is formed by admixture between two or more subpopulations, confounding can be estimated and controlled. Current implementations of this approach have limitations, the most serious of which is that they do not allow for uncertainty in estimations of individual admixture proportions or for lack of identifiability of subpopulations in the model. We describe methods that overcome these limitations by a combination of Bayesian and classical approaches, and we demonstrate the methods by using data from three admixed populations--African American, African Caribbean, and Hispanic American--in which there is extreme confounding of trait-genotype associations because the trait under study (skin pigmentation) varies with admixture proportions. In these data sets, as many as one-third of marker loci show crude associations with the trait. Control for confounding by population stratification eliminates these associations, except at loci that are linked to candidate genes for the trait. With only 32 markers informative for ancestry, the efficiency of the analysis is 70%. These methods can deal with both confounding and selection bias in genetic-association studies, making family-based designs unnecessary.

Bayes Theorem↗

A benchmark methodology for managing uncertainties in urban runoff quality models.

In this paper we present a benchmarking methodology, which aims at comparing urban runoff quality models, based on the Bayesian theory. After choosing the different configurations of models to be tested, this methodology uses the Metropolis algorithm, a general MCMC sampling method, to estimate the posterior distributions of the models' parameters. The analysis of these posterior distributions allows a quantitative assessment of the parameters' uncertainties and their interaction structure, and provides information about the sensitivity of the probability distribution of the model output to parameters. The effectiveness and efficiency of this methodology are illustrated in the context of 4 configurations of pollutants' accumulation/erosion models, tested on 4 street subcatchments. Calibration results demonstrate that the Metropolis algorithm produces reliable inferences of parameters thus, helping on the improvement of the mathematical concept of model equations.

Algorithms↗

Bayesian approaches to multiple sources of evidence and uncertainty in complex cost-effectiveness modelling.

Increasingly complex models are being used to evaluate the cost-effectiveness of medical interventions. We describe the multiple sources of uncertainty that are relevant to such models, and their relation to either probabilistic or deterministic sensitivity analysis. A Bayesian approach appears natural in this context. We explore how sensitivity analysis to patient heterogeneity and parameter uncertainty can be simultaneously investigated, and illustrate the necessary computation when expected costs and benefits can be calculated in closed form, such as in discrete-time discrete-state Markov models. Information about parameters can either be expressed as a prior distribution, or derived as a posterior distribution given a generalized synthesis of available data in which multiple sources of evidence can be differentially weighted according to their assumed quality. The resulting joint posterior distributions on costs and benefits can then provide inferences on incremental cost-effectiveness, best presented as posterior distributions over net-benefit and cost-effectiveness acceptability curves. These ideas are illustrated with a detailed running example concerning the cost-effectiveness of hip prostheses in different age-sex subgroups. All computations are carried out using freely available software for conducting Markov chain Monte Carlo analysis.

Adult↗

On the interpretation of certainty factors in expert systems.

Despite the strong theoretical foundation the Bayesian probabilistic approach to model uncertainty in medicine meets many difficulties at the implementation step. One of these difficulties is related to a large amount of conditional probabilities to be assessed and in many cases this task was recognised to be practically insoluble. The MYCIN certainty factors model is a widely distributed pragmatical approach for modeling reasoning under uncertainty that substantially simplifies the problem, at the sacrifice of theoretical soundness. One can determine certainty factors as a function of prior and posterior probability. However, this approach is only consistent with the modularity axiom for certainty factors for tree-structure inference networks, which is rarely true for practical applications. In this paper we abandon the requirement of a direct probabilistic interpretation of certainty factors and build a model of propagation of uncertainty in terms of absolute belief and belief updates. We describe our model for propagating uncertainty in terms of matrix multiplication with specifically defined addition and multiplication which correspond to parallel and sequential combinations of certainty factors. It is possible to define these operations in such a manner that they form a field, and therefore to obtain some useful properties. Finally we present a method of determining certainty factors from statistical data using nonlinear regression and illustrate it with a leukemia diagnostics problem.

Artificial Intelligence↗

Monophyly and relationships of wrens (Aves: Troglodytidae): a congruence analysis of heterogeneous mitochondrial and nuclear DNA sequence data.

The wrens (Aves: Troglodytidae) are a group of primarily New World insectivorous birds, the monophyly of which has long been recognized, but whose intergeneric relationships are essentially unknown. In order to test the monophyly of the group, and to attempt to resolve relationships among genera within it, sequences from the mitochondrial cytochrome b gene and the fourth intron of the nuclear beta-fibrinogen gene were obtained from nearly all genera of wrens, from their relatives as suggested by traditional taxonomy and DNA-DNA hybridization analyses, and from additional passerines. Maximum likelihood analysis of the two data sets yielded maximal congruence between independently derived estimates of relationship, outperforming a variety of weighted parsimony methods. Hierarchical likelihood ratio tests indicated that the two gene regions differed significantly in every estimated parameter of sequence evolution, and combined analysis of the two data sets was accomplished using a heterogeneous-model Bayesian approach. Independent and simultaneous analyses of both data sets supported monophyly of the wrens (excluding one recently added member, the monotypic genus Donacobius) and a sister-group relationship between wrens and the gnatcatchers (Polioptila). Additionally, strong support was found for paraphyly of the genus Thryothorus, and for a sister-group relationship between the genera Cistothorus and Troglodytes. Analyses of these data failed to resolve basal relationships within wrens, possibly due to ambiguity in rooting with a distant, species-poor outgroup. Analysis of the combined data for wrens alone yielded results which were largely congruent with relationships inferred using the complete data set, with the benefit of stronger support for relationships within the group. However, alternative rootings of this ingroup tree were weakly supported by nucleotide substitution data. Insertion-deletion events suggest that the genus Salpinctes may be sister to all other wrens.

Animals↗

Testing for differentiation of microbial communities using phylogenetic methods: accounting for uncertainty of phylogenetic inference and character state mapping.

Comparative analyses of microbial communities increasingly involve the assay of 16S rRNA (or other gene) sequences from environmental DNA. Determining whether the composition of two or more communities differ in their phylogenetic composition involves testing for covariation between phylogeny and community type. This approach requires estimating the phylogenetic relationships among all sampled sequences and assessing whether the distribution of sequences among communities differs from the null expectation that sequences are randomly distributed. One method developed for implementing the phylogeny-based test of differentiation, referred to as the Phylogenetic test, relies on a single estimate of the phylogeny. However, for most data sets, many alternative phylogenetic trees provide statistically equivalent descriptions of the data. Because the actual phylogeny is unknown, phylogenetic tests of differentiation among microbial communities must account for phylogenetic uncertainty. In this article, we evaluate bootstrapping and Bayesian phylogenetic methods when implementing the Phylogenetic test using parsimony to map character states, and we investigate the effects of character mapping uncertainty by using a Bayesian approach to stochastically map character states on trees. Our approaches incorporate uncertainty into the tests of two closely related null hypotheses: (1) populations are panmictic, and (2) identical communities existed in both environments over the course of evolutionary history. We use two data sets previously implemented in tests for community differentiation: nitrite reductase genes sampled from marsh and upland soils and 16S rDNA sequences sampled from the human mouth and gut. We show that accounting for phylogenetic and mapping uncertainties can drastically affect results when implementing the Phylogenetic test. Accounting for phylogenetic and character mapping uncertainty provides a more conservative and robust test of covariation between phylogeny and environment when comparing microbial communities using DNA sequences.

Bacteria↗

Bayesian second-level analysis of functional magnetic resonance images.

We propose a new method for the second-level analysis of functional MRI data based on Bayesian statistics. Our method does not require a computationally costly Bayesian model on the first level of analysis. Rather, modeling for single subjects is realized by means of the commonly applied General Linear Model. On the basis of the resulting parameter estimates for single subjects we calculate posterior probability maps and maps of the effect size for effects of interest in groups of subjects. A comparison of this method with the conventional analysis based on t statistics shows that the new approach is more robust against outliers. Moreover, our method overcomes some of the severe problems of null hypothesis significance tests such as the need to correct for multiple comparisons and facilitates inferences which are hard to formulate in terms of classical inferences.

Algorithms↗