Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

Host-symbiont stability and fast evolutionary rates in an ant-bacterium association: cospeciation of camponotus species and their endosymbionts, candidatus blochmannia.

Bacterial endosymbionts are widespread across several insect orders and are involved in interactions ranging from obligate mutualism to reproductive parasitism. Candidatus Blochmannia gen. nov. (Blochmannia) is an obligate bacterial associate of Camponotus and related ant genera (Hymenoptera: Formicidae). The occurrence of Blochmannia in all Camponotus species sampled from field populations and its maternal transmission to host offspring suggest that this bacterium is engaged in a long-term, stable association with its ant hosts. However, evidence for cospeciation in this system is equivocal because previous phylogenetic studies were based on limited gene sampling, lacked statistical analysis of congruence, and have even suggested host switching. We compared phylogenies of host genes (the nuclear EF-1alphaF2 and mitochondrial COI/II) and Blochmannia genes (16S ribosomal DNA [rDNA], groEL, gidA, and rpsB), totaling more than 7 kilobases for each of 16 Camponotus species. Each data set was analyzed using maximum likelihood and Bayesian phylogenetic reconstruction methods. We found minimal conflict among host and symbiont phylogenies, and the few areas of discordance occurred at deep nodes that were poorly supported by individual data sets. Concatenated protein-coding genes produced a very well-resolved tree that, based on the Shimodaira-Hasegawa test, did not conflict with any host or symbiont data set. Correlated rates of synonymous substitution (d(S)) along corresponding branches of host and symbiont phylogenies further supported the hypothesis of cospeciation. These findings indicate that Blochmannia-Camponotus symbiosis has been evolutionarily stable throughout tens of millions of years. Based on inferred divergence times among the ant hosts, we estimated rates of sequence evolution of Blochmannia to be approximately 0.0024 substitutions per site per million years (s/s/MY) for the 16S rDNA gene and approximately 0.1094 s/s/MY at synonymous positions of the genes sampled. These rates are several-fold higher than those for related bacteria Buchnera aphidicola and Escherichia coli. Phylogenetic congruence among Blochmannia genes indicates genome stability that typifies primary endosymbionts of insects.

Animals↗

Functional diversification of B MADS-box homeotic regulators of flower development: Adaptive evolution in protein-protein interaction domains after major gene duplication events.

B-class MADS-box genes have been shown to be the key regulators of petal and stamen specification in several eudicot model species such as Arabidopsis thaliana, Antirrhinum majus, and Petunia hybrida. Orthologs of these genes have been found across angiosperms and gymnosperms, and it is thought that the basic regulatory function of B proteins is conserved in seed plant lineages. The evolution of B genes is characterized by numerous duplications that might represent key elements fostering the functional diversification of duplicates with a deep impact on their role in the evolution of the floral developmental program. To evaluate this, we performed a rigorous statistical analysis with B gene sequences. Using maximum likelihood and Bayesian methods, we estimated molecular substitution rates and determined the selective regimes operating at each residue of B proteins. We implemented tests that rely on phylogenetic hypotheses and codon substitution models to detect significant differences in substitution rates (DSRs) and sites under positive adaptive selection (PS) in specific lineages before and after duplication events. With these methods, we identified several protein residues fixed by PS shortly after the origin of PISTILLATA-like and APETALA3-like lineages in angiosperms and shortly after the origin of the euAP3-like lineage in core eudicots, the 2 main B gene duplications. The residues inferred to have been fixed by positive selection lie mostly within the K domain of the protein, which is key to promote heterodimerization. Additionally, we used a likelihood method that accommodates DSRs among lineages to estimate duplication dates for AP3-PI and euAP3-TM6, calibrating with data from the fossil record. The dates obtained are consistent with angiosperm origins and diversification of core eudicots. Our results strongly suggest that novel multimer formation with other MADS proteins could have been crucial for the functional divergence of B MADS-box genes. We thus propose a mechanism of functional diversification and persistence of gene duplicates by the appearance of novel multimerization capabilities after duplications. Multimer formation in different combinations of regulatory proteins can be a mechanistic basis for the origin of novel regulatory functions and a gene regulatory mechanism for the appearance of morphological innovations.

Amino Acid Sequence↗

The effect of ignoring individual heterogeneity in Weibull log-normal sire frailty models.

The objective of this study was, by means of simulation, to quantify the effect of ignoring individual heterogeneity in Weibull sire frailty models on parameter estimates and to address the consequences for genetic inferences. Three simulation studies were evaluated, which included 3 levels of individual heterogeneity combined with 4 levels of censoring (0, 25, 50, or 75%). Data were simulated according to balanced half-sib designs using Weibull log-normal animal frailty models with a normally distributed residual effect on the log-frailty scale. The 12 data sets were analyzed with 2 models: the sire model, equivalent to the animal model used to generate the data (complete sire model), and a corresponding model in which individual heterogeneity in log-frailty was neglected (incomplete sire model). Parameter estimates were obtained from a Bayesian analysis using Gibbs sampling, and also from the software Survival Kit for the incomplete sire model. For the incomplete sire model, the Monte Carlo and Survival Kit parameter estimates were similar. This study established that when unobserved individual heterogeneity was ignored, the parameter estimates that included sire effects were biased toward zero by an amount that depended in magnitude on the level of censoring and the size of the ignored individual heterogeneity. Despite the biased parameter estimates, the ranking of sires, measured by the rank correlations between true and estimated sire effects, was unaffected. In comparison, parameter estimates obtained using complete sire models were consistent with the true values used to simulate the data. Thus, in this study, several issues of concern were demonstrated for the incomplete sire model.

Animals↗

Beyond statistical inference: a decision theory for science.

Traditional null hypothesis significance testing does not yield the probability of the null or its alternative and, therefore, cannot logically ground scientific decisions. The decision theory proposed here calculates the expected utility of an effect on the basis of (1) the probability of replicating it and (2) a utility function on its size. It takes significance tests--which place all value on the replicability of an effect and none on its magnitude--as a special case, one in which the cost of a false positive is revealed to be an order of magnitude greater than the value of a true positive. More realistic utility functions credit both replicability and effect size, integrating them for a single index of merit. The analysis incorporates opportunity cost and is consistent with alternate measures of effect size, such as r2 and information transmission, and with Bayesian model selection criteria. An alternate formulation is functionally equivalent to the formal theory, transparent, and easy to compute.

Data Interpretation, Statistical↗

Empirical bayes estimation of a sparse vector of gene expression changes.

Gene microarray technology is often used to compare the expression of thousand of genes in two different cell lines. Typically, one does not expect measurable changes in transcription amounts for a large number of genes; furthermore, the noise level of array experiments is rather high in relation to the available number of replicates. For the purpose of statistical analysis, inference on the "population'' difference in expression for genes across the two cell lines is often cast in the framework of hypothesis testing, with the null hypothesis being no change in expression. Given that thousands of genes are investigated at the same time, this requires some multiple comparison correction procedure to be in place. We argue that hypothesis testing, with its emphasis on type I error and family analogues, may not address the exploratory nature of most microarray experiments. We instead propose viewing the problem as one of estimation of a vector known to have a large number of zero components. In a Bayesian framework, we describe the prior knowledge on expression changes using mixture priors that incorporate a mass at zero, and we choose a loss function that favors the selection of sparse solutions. We consider two different models applicable to the microarray problem, depending on the nature of replicates available, and show how to explore the posterior distributions of the parameters using MCMC. Simulations show an interesting connection between this Bayesian estimation framework and false discovery rate (FDR) control. Finally, two empirical examples illustrate the practical advantages of this Bayesian estimation paradigm.

Journal Article↗

Detecting correlation between characters in a comparative analysis with uncertain phylogeny.

The importance of accommodating the phylogenetic history of a group when performing a comparative analysis is now widely recognized. The typical approaches either assume the tree is known without error, or they base inferences on a collection of well-supported trees or on a collection of trees generated under a stochastic model of cladogenesis. However, these approaches do not adequately account for the uncertainty of phylogenetic trees in a comparative analysis, especially when data relevant to the phylogeny of a group are available. Here, we develop a method for performing comparative analyses that is based on an extension of Felsenstein's independent contrasts method. Uncertainties in the phylogeny, branch lengths, and other parameters are accommodated by averaging over all possible trees, weighting each by the probability that the tree is correct. We do this in a Bayesian framework and use Markov chain Monte Carlo to perform the high-dimensional summations and integrations required by the analysis. We illustrate the method using comparative characters sampled from Anolis lizards.

Animals↗

Coalescent-based association mapping and fine mapping of complex trait loci.

We outline a general coalescent framework for using genotype data in linkage disequilibrium-based mapping studies. Our approach unifies two main goals of gene mapping that have generally been treated separately in the past: detecting association (i.e., significance testing) and estimating the location of the causative variation. To tackle the problem, we separate the inference into two stages. First, we use Markov chain Monte Carlo to sample from the posterior distribution of coalescent genealogies of all the sampled chromosomes without regard to phenotype. Then, averaging across genealogies, we estimate the likelihood of the phenotype data under various models for mutation and penetrance at an unobserved disease locus. The essential signal that these models look for is that in the presence of disease susceptibility variants in a region, there is nonrandom clustering of the chromosomes on the tree according to phenotype. The extent of nonrandom clustering is captured by the likelihood and can be used to construct significance tests or Bayesian posterior distributions for location. A novelty of our framework is that it can naturally accommodate quantitative data. We describe applications of the method to simulated data and to data from a Mendelian locus (CFTR, responsible for cystic fibrosis) and from a proposed complex trait locus (calpain-10, implicated in type 2 diabetes).

Alleles↗

Top-down influence in early visual processing: a Bayesian perspective.

Traditional views of visual processing suggest that early visual neurons are static spatiotemporal filters that extract local features by feedforward computation. The extracted information is then fed forward through a chain of modules to successively higher visual areas for further analysis. Recording from early visual neurons in awake behaving monkeys, we revealed there are many levels of complexity in the information processing of the early visual cortex. We found that the early visual neurons not only are sensitive to features within their receptive fields (RFs) but also to the global context of a visual scene, the behavioral relevance of the stimuli and the experience of the animals. These findings suggest that the early visual cortex (V1 and V2) is tightly coupled to and highly interactive with the rest of the visual system. The top-down interaction, mediated by recurrent feedback connections, introduces contextual information to influence the perceptual inference in the early visual cortex.

Algorithms↗

Cost-effectiveness analysis when the WTA is greater than the WTP.

The incremental cost effectiveness ratio has long been the standard parameter of interest in the assessment of the cost-effectiveness of a new treatment. However, due to concerns with interpretability and statistical inference, authors have suggested using the willingness-to-pay for a unit of health benefit to define the incremental net benefit as an alternative. The incremental net benefit has a more consistent interpretation and is amenable to routine statistical procedures. These procedures rely on the fact that the willingness-to-accept compensation for a loss of a unit of health benefit (at some cost saving) is the same as the willingness-to-pay for it. Theoretical and empirical evidence suggest, however, that in health care the willingness-to-accept is about twice as much as the willingness-to-pay. We use Bayesian methods to provide a statistical procedure for the cost-effectiveness comparison of two arms of a randomized clinical trial that allows the willingness-to-pay and the willingness-to-accept to have different values. An example is provided.

Bayes Theorem↗

A future for models and data in environmental science.

Together, graphical models and the Bayesian paradigm provide powerful new tools that promise to change the way that environmental science is done. The capacity to merge theory with mechanistic understanding and empirical evidence, to assimilate diverse sources of information and to accommodate complexity will transform the collection and interpretation of data. As we discuss here, we specifically expect a shift from a focus on simple experiments with inflexible design and selection among models that embrace parts of processes to a synthesis of integrated process models. With this potential come new challenges, including some that are specific and technical and others that are general and will involve reexamination of the role of inference and prediction.

Animals↗

Haplotype reconstruction for diploid populations.

The inference of haplotype pairs directly from unphased genotype data is a key step in the analysis of genetic variation in relation to disease and pharmacogenetically relevant traits. Most popular methods such as Phase and PL do require either the coalescence assumption or the assumption of linkage between the single-nucleotide polymorphisms (SNPs). We have now developed novel approaches that are independent of these assumptions. First, we introduce a new optimization criterion in combination with a block-wise evolutionary Monte Carlo algorithm. Based on this criterion, the 'haplotype likelihood', we develop two kinds of estimators, the maximum haplotype-likelihood (MHL) estimator and its empirical Bayesian (EB) version. Using both real and simulated data sets, we demonstrate that our proposed estimators allow substantial improvements over both the expectation-maximization (EM) algorithm and Clark's procedure in terms of capacity/scalability and error rate. Thus, hundreds and more ambiguous loci and potentially very large sample sizes can be processed. Moreover, applying our proposed EB estimator can result in significant reductions of error rate in the case of unlinked or only weakly linked SNPs.

Algorithms↗

Numbers of mutations to different types of colorectal cancer.

BACKGROUND: The numbers of oncogenic mutations required for transformation are uncertain but may be inferred from how cancer frequencies increase with aging. Cancers requiring more mutations will tend to appear later in life. This type of approach may be confounded by biologic heterogeneity because different cancer subtypes may require different numbers of mutations. For example, a sporadic cancer should require at least one more somatic mutation relative to its hereditary counterpart. METHODS: To better estimate numbers of mutations before transformation, 1,022 colorectal cancers were classified with respect to microsatellite instability (MSI) and germline DNA mismatch repair mutations characteristic of hereditary nonpolyposis colorectal cancer (HNPCC). MSI- cancers were also classified with respect to clinical stage. Ages at cancer and a Bayesian algorithm were used to estimate the numbers of oncogenic mutations required for transformation for each cancer subtype. RESULTS: Ages at MSI+ cancers were consistent with five or six oncogenic mutations for hereditary (HNPCC) cancers, and seven or eight mutations for its sporadic counterpart. Ages at cancer were consistent with seven mutations for sporadic MSI- cancers, and were similar (six to eight mutations) regardless of clinical cancer stage. CONCLUSION: Different biologic subtypes of colorectal cancer appear to require different numbers of oncogenic mutations before transformation. Sporadic MSI+ cancers may require more than a single additional somatic alteration compared to hereditary MSI+ cancers because the epigenetic inactivation of MLH1 commonly observed in sporadic MSI+ cancers may be a multistep process. Interestingly, estimated numbers of MSI- cancer mutations were similar (six to eight mutations) regardless of clinical cancer stage, suggesting a propensity to spread or metastasize does not require additional mutations after transformation. Estimates of oncogenic mutation numbers may help explain some of the biology underlying different cancer subtypes.

Adult↗

Phylogeny and toxigenic potential is correlated in Fusarium species as revealed by partial translation elongation factor 1 alpha gene sequences.

Partial translation elongation factor 1 alpha (TEF-1alpha) gene and intron sequences are reported from 148 isolates of 11 species of the anamorph genus Fusarium; F. avenaceum (syn. F. arthrosporioides), F. cerealis, F. culmorum, F. equiseti, F.flocciferum, F. graminearum, F. lunulosporum, F. sambucinum, F. torulosum, F. tricinctum and F. venenatum. The sequences were aligned with TEF-1alpha sequences retrieved from 35 isolates of F. kyushuense, F. langsethiae, F. poae and F. sporotrichioides in a previous study, and 39 isolates of F. cerealis, F. culmorum, F. graminearum and F. pseudograminearum retrieved from sequence databases. The 222 aligned sequences were subjected to phylogenetic analyses using maximum parsimony and Bayesian Markov Chain Monte Carlo maximum likelihood statistics. Support for internal branching topologies was examined by Bremer support, bootstrap and posterior probability analyses. The resulting trees were largely congruent. The taxon groups included in the sections Discolor, Gibbosum and Sporotrichiella sensu Wollenweber & Reinking (1935) all appeared to be polyphyletic. All species were monophyletic except F. flocciferum that was paraphyletic, and one isolate classified as F. cfr langsethiae on the basis of morphology that grouped with F. sporotrichioides. Mapping of toxin profiles, host preferences and geographic origin onto the DNA based phylogenetic tree structure indicated that in particular the toxin profiles corresponded with phylogeny, i.e. phylotoxigenic relationships were inferred. A major distinction was observed between the trichothecene and non-trichothecene producers, and the trichothecene producers were grouped into one clade of strictly type A trichothecene producers, one clade of strictly type B trichothecene producers and one clade with both type A and type B trichothecene producers. Furthermore, production of the type A trichothecenes T-2/HT-2 toxins are associated with a lineage comprising F. langsethiae and F. sporotrichioides. The ability to produce zearalenone was apparently gained parallel to the ability to produce trichothecenes, and later lost in a derived sublineage. The ability to produce enniatins is a shared feature of the entire study group, with the exception of the strict trichothecene type B producers and F. equiseti. The ability to produce moniliformin seems to be an ancestral feature of members of the genus Fusarium which seems to have been lost in the clades consisting of trichothecene/zearalenone producers. The aims of the present study were to determine the phylogenetic relationships between the different species of Fusarium commonly occurring on Norwegian cereals and some of their closest relatives, as well as to reveal underlying patterns such as the ability to produce certain mycotoxins, geographic distribution and host preferences. Implications for a better classification of Fusarium are discussed and highlighted.

DNA, Fungal↗

Geostatistical analysis of disease data: accounting for spatial support and population density in the isopleth mapping of cancer mortality risk using area-to-point Poisson kriging.

BACKGROUND: Geostatistical techniques that account for spatially varying population sizes and spatial patterns in the filtering of choropleth maps of cancer mortality were recently developed. Their implementation was facilitated by the initial assumption that all geographical units are the same size and shape, which allowed the use of geographic centroids in semivariogram estimation and kriging. Another implicit assumption was that the population at risk is uniformly distributed within each unit. This paper presents a generalization of Poisson kriging whereby the size and shape of administrative units, as well as the population density, is incorporated into the filtering of noisy mortality rates and the creation of isopleth risk maps. An innovative procedure to infer the point-support semivariogram of the risk from aggregated rates (i.e. areal data) is also proposed. RESULTS: The novel methodology is applied to age-adjusted lung and cervix cancer mortality rates recorded for white females in two contrasted county geographies: 1) state of Indiana that consists of 92 counties of fairly similar size and shape, and 2) four states in the Western US (Arizona, California, Nevada and Utah) forming a set of 118 counties that are vastly different geographical units. Area-to-point (ATP) Poisson kriging produces risk surfaces that are less smooth than the maps created by a naïve point kriging of empirical Bayesian smoothed rates. The coherence constraint of ATP kriging also ensures that the population-weighted average of risk estimates within each geographical unit equals the areal data for this unit. Simulation studies showed that the new approach yields more accurate predictions and confidence intervals than point kriging of areal data where all counties are simply collapsed into their respective polygon centroids. Its benefit over point kriging increases as the county geography becomes more heterogeneous. CONCLUSION: A major limitation of choropleth maps is the common biased visual perception that larger rural and sparsely populated areas are of greater importance. The approach presented in this paper allows the continuous mapping of mortality risk, while accounting locally for population density and areal data through the coherence constraint. This form of Poisson kriging will facilitate the analysis of relationships between health data and putative covariates that are typically measured over different spatial supports.

Cluster Analysis↗

Bayesian modeling of animal- and herd-level prevalences.

We reviewed Bayesian approaches for animal-level and herd-level prevalence estimation based on cross-sectional sampling designs and demonstrated fitting of these models using the WinBUGS software. We considered estimation of infection prevalence based on use of a single diagnostic test applied to a single herd with binomial and hypergeometric sampling. We then considered multiple herds under binomial sampling with the primary goal of estimating the prevalence distribution and the proportion of infected herds. A new model is presented that can be used to estimate the herd-level prevalence in a region, including the posterior probability that all herds are non-infected. Using this model, inferences for the distribution of prevalences, mean prevalence in the region, and predicted prevalence of herds in the region (including the predicted probability of zero prevalence) are also available. In the models presented, both animal- and herd-level prevalences are modeled as mixture distributions to allow for zero infection prevalences. (If mixture models for the prevalences were not used, prevalence estimates might be artificially inflated, especially in herds and regions with low or zero prevalence.) Finally, we considered estimation of animal-level prevalence based on pooled samples.

Animals↗

Molecular phylogenetics of the clover genus (Trifolium--Leguminosae).

Trifolium, the clover genus, is one of the largest genera of the legume family. We conducted parsimony and Bayesian phylogenetic analyses based on nuclear ribosomal DNA internal transcribed spacer and chloroplast trnL intron sequences obtained from 218 of the ca. 255 species of Trifolium, representatives from 11 genera of the vicioid clade, and an outgroup Lotus. We confirm the monophyly of Trifolium, and propose a new infrageneric classification of the genus based on the phylogenetic results. Incongruence between the nrDNA and cpDNA results suggests five to six cases of apparent hybrid speciation, and identifies the putative progenitors of the allopolyploids T. dubium, a widespread weed, and T. repens, the most commonly cultivated clover species. Character state reconstructions confirm 2n=16 as the ancestral chromosome number in Trifolium, and infer a minimum of 19 instances of aneuploidy and 22 of polyploidy in the genus. The ancestral life history is hypothesized to be annual in subgenus Chronosemium and equivocal in subgenus Trifolium. Transitions between the annual and perennial habit are common. Our results are consistent with a Mediterranean origin of the genus, probably in the Early Miocene. A single origin of all North and South American species is hypothesized, while the species of sub-Saharan Africa may originate from three separate dispersal events.

Bayes Theorem↗

Mitochondrial DNA phylogeography of Lissotriton boscai (Caudata, Salamandridae): evidence for old, multiple refugia in an Iberian endemic.

In Europe, southern peninsulas served as refugia during cold periods in the Pleistocene, acting both as centres of origin of endemisms and as sources from which formerly glaciated areas were recolonized during interglacial periods. Previous studies have revealed that within the main refugial areas, intraspecific lineages often survived in allopatric refugia. We analysed two mitochondrial markers (nad4, control region, approximately 1.4 kb) in 103 individuals representing the entire distribution of Lissotriton boscai, a newt endemic to the western Iberian Peninsula. We inferred the evolutionary history of the species through phylogenetic, phylogeographic and historical demographic analyses. The results revealed unexpected, deep levels of geographically structured genetic variability. We identified two main evolutionary lineages, each containing three well-supported clades. The first historical split involved populations from central-southwestern coastal Portugal and the ancestor of all the remaining populations around 5.8 million years ago. Both lineages were subsequently fragmented into different population groups between 2.5 and 1.2 million years ago. According to nested clade analysis, at lower hierarchical levels the patterns suggest restricted gene flow with isolation by distance, whereas at higher levels the clades exhibit signatures of contiguous range expansion. Bayesian Skyline Plots show recent bottlenecks, followed by demographic expansions in all lineages. The significant genetic structure found is consistent with long-term survival of populations in allopatric refugia, supporting the 'refugia-within-refugia' scenario for southern European peninsulas. The comparison of our results with other co-distributed species highlights the generality of this hypothesis for the Iberian herpetofauna and suggests that Mediterranean refuges had more relevance for the composition and distribution of present biodiversity patterns than currently acknowledged. We briefly discuss the taxonomic and conservation implications of our results.

Animals↗

Admixed sexual and facultatively asexual aphid lineages at mating sites.

Cyclically parthenogenetic organisms may have facultative asexual counterparts. Such organisms, including aphids, are therefore interesting models for the study of ecological and genetic interactions between lineages differing in reproductive mode. Earlier studies on aphids have revealed major differences in the genetic outcomes of populations that are possibly resulting mostly either from sexual or from asexual reproduction. Besides, notable gene flow between sexual and asexual derivatives has been suspected, which could lead to the emergence of new asexual lineages. The present study examines the interplay between these lineages and is based on analyses of population structure of individuals that may contribute to the pool of sexual reproductive forms in the host alternating aphid Rhopalosiphum padi. Using a Bayesian assignment method, we first show that the sexual forms of R. padi on mating sites encompass two genetically distinct clusters of individuals in the western part of France. The first cluster included unique genotypes of sexual lineages, while the second cluster included facultatively asexual lineages in numerous copies, the reproductive mode of the two clusters being confirmed by reference clones. Sexual reproductive forms produced by sexual and facultatively asexual lineages are thus admixed at mating sites which gives a large opportunity for the two clusters to mate with each other. Nevertheless, this study also highlights, as previously demonstrated, that the two clusters retained high genetic differentiation. Possible explanations for the inferred limited genetic exchanges are advanced in the discussion, but further dedicated investigations are required to solve this paradox.

Animals↗