Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

Clinical inferences and decisions--II. Decision trees, receiver operator curves and subjective probability.

In patient management, clinical decisions follow a logical sequence which can be formally expressed as a decision tree in which the uncertainties associated with each alternative outcome may be made explicit using Bayes' theorem. Where test data is used in the formulation of a decision, the uncertainty associated with the information it conveys may be modified by changing the pass/fail criterion to alter the false positive and false negative error rate. Classical procedures based on information theory are described to illustrate how this may be achieved for any test. When hard data is not available to permit such an approach, the clinician must rely on his own past experience or that of a colleague. Several methods are available for quantifying such experience by estimating subjective probabilities associated with an action or test result. Two simple methods are described for deriving subjective probabilities for subsequent use within a Bayesian decision model.

Bayes Theorem↗

Dracula ant phylogeny as inferred by nuclear 28S rDNA sequences and implications for ant systematics (Hymenoptera: Formicidae: Amblyoponinae).

Ants are one of the most ecologically and numerically dominant families of organisms in almost every terrestrial habitat throughout the world, though they include only about 1% of all described insect species. The development of eusociality is thought to have been a driving force in the striking diversification and dominance of this group, yet we know little about the evolution of the major lineages of ants and have been unable to clearly determine their primitive characteristics. Ants within the subfamily Amblyoponinae are specialized arthropod predators, possess many anatomically and behaviorally primitive characters and have been proposed as a possible basal lineage within the ants. We investigate the phylogenetic relationships among the members of the subfamily, using nuclear 28S rDNA sequence data. Outgroups for the analysis include members of the poneromorph and leptanillomorph (Apomyrma, Leptanilla) ant subfamilies, as well as three wasp families. Parsimony, maximum likelihood, and Bayesian analyses provide strong support for the monophyly of a clade containing the two genera Apomyrma+Mystrium (100% bpp; 97% ML bs; and 97% MP bs), and moderate support for the monophyly of the Amblyoponinae as long as Apomyrma (Apomyrminae) is included (87% bpp; 57% ML bs; and 76% MP bs). Analyses did not recover evidence of monophyly of the Amblyopone genus, while the monophyly of the other genera in the subfamily is supported. Based on these results we provide a morphological diagnosis of the Amblyoponinae that includes Apomyrma. Among the outgroup taxa, Typhlomyrmex grouped consistently with Ectatomma, supporting the recent placement of Typhlomyrmex in the Ectatomminae. The results of this present study place the included ant subfamilies into roughly two clades with the basal placement of Leptanilla unclear. One clade contains all the Amblyoponinae (including Apomyrma), Ponerinae, and Proceratiinae (Poneroid clade). The other clade contains members from subfamilies Cerapachyinae, Dolichoderinae, Ectatomminae, Formicinae, Myrmeciinae, and Myrmicinae (Formicoid clade).

Animals↗

Temporal reasoning for diagnosis in a causal probabilistic knowledge base.

We have added temporal reasoning to the Heart Disease Program (HDP) to take advantage of the temporal constraints inherent in cardiovascular reasoning. Some processes take place over minutes while others take place over months or years and a strictly probabilistic formalism can generate hypotheses that are impossible given the temporal relationships involved. The HDP has temporal constraints on the causal relations specified in the knowledge base and temporal properties on the patient input provided by the user. These are used in two ways. First, they are used to constrain the generation of the pre-computed causal pathways through the model that speed the generation of hypotheses. Second, they are used to generate time intervals for the instantiated nodes in the hypotheses, which are matched and adjusted as nodes are added to each evolving hypothesis. This domain offers a number of challenges for temporal reasoning. Since the nature of diagnostic reasoning is inferring a causal explanation from the evidence, many of the temporal intervals have few constraints and the reasoning has to make maximum use of those that exist. Thus, the HDP uses a temporal interval representation that includes the earliest and latest beginning and ending specified by the constraints. Some of the disease states can be corrected but some of the manifestations may remain. For example, a valve disease such as aortic stenosis produces hypertrophy that remains long after the valve has been replaced. This requires multiple time intervals to account for the existing findings. This paper discusses the issues and solutions that have been developed for temporal reasoning integrated with a pseudo-Bayesian probabilistic network in this challenging domain for diagnosis.

Artificial Intelligence↗

Large hierarchical Bayesian analysis of multivariate survival data.

Failure times that are grouped according to shared environments arise commonly in statistical practice. That is, multiple responses may be observed for each of many units. For instance, the units might be patients or centers in a clinical trial setting. Bayesian hierarchical models are appropriate for data analysis in this context. At the first stage of the model, survival times can be modelled via the Cox partial likelihood, using a justification due to Kalbfleisch (1978, Journal of the Royal Statistical Society, Series B 40, 214-221). Thus, questionable parametric assumptions are avoided. Conventional wisdom dictates that it is comparatively safe to make parametric assumptions at subsequent stages. Thus, unit-specific parameters are modelled parametrically. The posterior distribution of parameters given observed data is examined using Markov chain Monte Carlo methods. Specifically, the hybrid Monte Carlo method, as described by Neal (1993a, in Advances in Neural Information Processing 5, 475-482; 1993b, Probabilistic inference using Markov chain Monte Carlo methods), is utilized.

Antineoplastic Combined Chemotherapy Protocols↗

A Bayesian evolutionary distance for parametrically aligned sequences.

There is an inherent relationship between the process of pairwise sequence alignment and the estimation of evolutionary distance. This relationship is explored and made explicit. Assuming an evolutionary model and given a specific pattern of observed base mismatches, the relative probabilities of evolution at each evolutionary distance are computed using a Bayesian framework. The mean or the median of this probability distribution provides a robust estimate of the central value. The evolutionary distance has traditionally been computed as zero for an observed homology of 20 bases with no mismatches; we prove that it is highly probable that the distance is greater than 0.01. The mean of the distribution is 0.047, which is a better estimate of the evolutionary distance. Bayesian estimates of the evolutionary distance incorporate arbitrary prior information about variable mutation rates both over time and along sequence position, thus requiring only a weak form of the molecular-clock hypothesis. The endpoints of the similarity between genomic DNA sequences are often ambiguous. The probability of evolution at each evolutionary distance can be estimated over the entire set of alignments by choosing the best alignment at each distance and the corresponding probability of duplication at that evolutionary distance. A central value of this distribution provides a robust evolutionary distance estimate. We provide an efficient algorithm for computing the parametric alignment, considering evolutionary distance as the only parameter. These techniques and estimates are used to infer the duplication history of the genomic sequence in C. elegans and in S. cerevisiae. Our results indicate that repeats discovered using a single scoring matrix show a considerable bias in subsequent evolutionary distance estimates.

Animals↗

Problems due to small samples and sparse data in conditional logistic regression analysis.

Conditional logistic regression was developed to avoid "sparse-data" biases that can arise in ordinary logistic regression analysis. Nonetheless, it is a large-sample method that can exhibit considerable bias when certain types of matched sets are infrequent or when the model contains too many parameters. Sparse-data bias can cause misleading inferences about confounding, effect modification, dose response, and induction periods, and can interact with other biases. In this paper, the authors describe these problems in the context of matched case-control analysis and provide examples from a study of electrical wiring and childhood leukemia and a study of diet and glioma. The same problems can arise in any likelihood-based analysis, including ordinary logistic regression. The problems can be detected by careful inspection of data and by examining the sensitivity of estimates to category boundaries, variables in the model, and transformations of those variables. One can also apply various bias corrections or turn to methods less sensitive to sparse data than conditional likelihood, such as Bayesian and empirical-Bayes (hierarchical regression) methods.

Bias↗

Population genetics and phylogenetic inference in bacterial molecular systematics: the roles of migration and recombination in Bradyrhizobium species cohesion and delineation.

A combination of population genetics and phylogenetic inference methods was used to delineate Bradyrhizobium species and to uncover the evolutionary forces acting at the population-species interface of this bacterial genus. Maximum-likelihood gene trees for atpD, glnII, recA, and nifH loci were estimated for diverse strains from all but one of the named Bradyrhizobium species, and three unnamed "genospecies," including photosynthetic isolates. Topological congruence and split decomposition analyses of the three housekeeping loci are consistent with a model of frequent homologous recombination within but not across lineages, whereas strong evidence was found for the consistent lateral gene transfer across lineages of the symbiotic (auxiliary) nifH locus, which grouped strains according to their hosts and not by their species assignation. A well resolved Bayesian species phylogeny was estimated from partially congruent glnII+recA sequences, which is highly consistent with the actual taxonomic scheme of the genus. Population-level analyses of isolates from endemic Canarian genistoid legumes based on REP-PCR genomic fingerprints, allozyme and DNA polymorphism analyses revealed a non-clonal and slightly epidemic population structure for B. canariense isolates of Canarian and Moroccan origin, uncovered recombination and migration as significant evolutionary forces providing the species with internal cohesiveness, and demonstrated its significant genetic differentiation from B. japonicum, its sister species, despite their sympatry and partially overlapped ecological niches. This finding provides strong evidence for the existence of well delineated species in the bacterial world. The results and approaches used herein are discussed in the context of bacterial species concepts and the evolutionary ecology of (brady)rhizobia.

Base Sequence↗

Probabilistic independent component analysis for functional magnetic resonance imaging.

We present an integrated approach to probabilistic independent component analysis (ICA) for functional MRI (FMRI) data that allows for nonsquare mixing in the presence of Gaussian noise. In order to avoid overfitting, we employ objective estimation of the amount of Gaussian noise through Bayesian analysis of the true dimensionality of the data, i.e., the number of activation and non-Gaussian noise sources. This enables us to carry out probabilistic modeling and achieves an asymptotically unique decomposition of the data. It reduces problems of interpretation, as each final independent component is now much more likely to be due to only one physical or physiological process. We also describe other improvements to standard ICA, such as temporal prewhitening and variance normalization of timeseries, the latter being particularly useful in the context of dimensionality reduction when weak activation is present. We discuss the use of prior information about the spatiotemporal nature of the source processes, and an alternative-hypothesis testing approach for inference, using Gaussian mixture models. The performance of our approach is illustrated and evaluated on real and artificial FMRI data, and compared to the spatio-temporal accuracy of results obtained from classical ICA and GLM analyses.

Algorithms↗

A multicompartment vascular model for inferring baseline and functional changes in cerebral oxygen metabolism and arterial dilation.

Functional hemodynamic responses are the composite results of underlying variations in cerebral oxygen consumption and the dilation of arterial vessels after neuronal activity. The development of biophysically based models of the cerebral vasculature allows the separation of the neuro-metabolic and neuro-vascular influences on measurable hemodynamic signals such as functional magnetic resonance imaging or optical imaging. We describe a multicompartment model of the vascular and oxygen transport dynamics associated with stimulus-driven neuronal activation. Our model offers several unique features compared with previous formulations such as the ability to estimate baseline blood flow, volume, and oxygen consumption from functional data. In addition, we introduce a capillary compliance model, arterial and venous oxygen permeability, and model the dynamics of extravascular tissue oxygenation. We apply this model to multimodal optical spectroscopic and laser speckle imaging of the rat somato-sensory cortex during nine conditions of whisker stimulation. By fitting the model using a psuedo-Bayesian framework to incorporate multimodal observations, we estimate baseline blood flow to be 94 (+/-15) mL/100 g min and baseline oxygen consumption to be 6.7 (+/-1.3) mL O(2)/100 g min. We calculate parametric, linear increases in arterial dilation (R(2)=0.96) and CMRO(2) (R(2)=0.87) responses over the nine conditions. Other parameters estimated by the model include vascular transit time and volume reserve, oxygen content, saturation, diffusivity rate constants, and partial pressure of oxygen in the vascular compartments and in the extravascular tissue. Finally, we compare this model to earlier work and find that the multicompartment model more accurately describes the observed oxygenation changes when compared with a single compartment version.

Arteries↗

Probe-level measurement error improves accuracy in detecting differential gene expression.

MOTIVATION: Finding differentially expressed genes is a fundamental objective of a microarray experiment. Numerous methods have been proposed to perform this task. Existing methods are based on point estimates of gene expression level obtained from each microarray experiment. This approach discards potentially useful information about measurement error that can be obtained from an appropriate probe-level analysis. Probabilistic probe-level models can be used to measure gene expression and also provide a level of uncertainty in this measurement. This probe-level measurement error provides useful information which can help in the identification of differentially expressed genes. RESULTS: We propose a Bayesian method to include probe-level measurement error into the detection of differentially expressed genes from replicated experiments. A variational approximation is used for efficient parameter estimation. We compare this approximation with MAP and MCMC parameter estimation in terms of computational efficiency and accuracy. The method is used to calculate the probability of positive log-ratio (PPLR) of expression levels between conditions. Using the measurements from a recently developed Affymetrix probe-level model, multi-mgMOS, we test PPLR on a spike-in dataset and a mouse time-course dataset. Results show that the inclusion of probe-level measurement error improves accuracy in detecting differential gene expression. AVAILABILITY: The MAP approximation and variational inference described in this paper have been implemented in an R package pplr. The MCMC method is implemented in Matlab. Both software are available from http://umber.sbs.man.ac.uk/resources/puma.

Algorithms↗

Combined nuclear and mitochondrial DNA sequences resolve generic relationships within the Cracidae (Galliformes, Aves).

The Cracidae is one of the most endangered and distinctive bird families in the Neotropics, yet the higher relationships among taxa remain uncertain. The molecular phylogeny of its 11 genera was inferred using 10,678 analyzable sites (5,412 from seven different mitochondrial segments and 5,266 sites from four nuclear genes). We performed combinability tests to check conflicts in phylogenetic signals of separate genes and genomes. Phylogenetic analysis showed that the unrooted tree of ((curassows, horned guan) (guans, chachalacas)) was favored by most data partitions and that different data partitions provided support for different parts of the tree. In particular, the concatenated mitochondrial DNA (mtDNA) genes resolved shallower nodes, whereas the combined nuclear sequences resolved the basal connections among the major clades of curassows, horned guan, chachalacas, and guans. Therefore, we decided that for the Cracidae all data should be combined for phylogenetic analysis. Maximum parsimony (MP), maximum likelihood (ML), and Bayesian analyses of this large data set produced similar trees. The MP tree indicated that guans are the sister group to (horned guan, (curassows, chachalacas)), whereas the ML and Bayesian analysis recovered a tree where the horned guan is a sister clade to curassows, and these two clades had the chachalacas as a sister group. Parametric bootstrapping showed that alternative trees previously proposed for the cracid genera are significantly less likely than our estimate of their relationships. A likelihood ratio test of the hypothesis of a molecular clock for cracid mtDNA sequences using the optimal ML topology did not reject rate constancy of substitutions through time. We estimated cracids to have originated between 64 and 90 million years ago (MYA), with a mean estimate of 76 MYA. Diversification of the genera occurred approximately 41-3 MYA, corresponding with periods of global climate change and other Earth history events that likely promoted divergences of higher level taxa.

Animals↗

Redundancy reduction revisited.

Soon after Shannon defined the concept of redundancy it was suggested that it gave insight into mechanisms of sensory processing, perception, intelligence and inference. Can we now judge whether there is anything in this idea, and can we see where it should direct our thinking? This paper argues that the original hypothesis was wrong in over-emphasizing the role of compressive coding and economy in neuron numbers, but right in drawing attention to the importance of redundancy. Furthermore there is a clear direction in which it now points, namely to the overwhelming importance of probabilities and statistics in neuroscience. The brain has to decide upon actions in a competitive, chance-driven world, and to do this well it must know about and exploit the non-random probabilities and interdependences of objects and events signalled by sensory messages. These are particularly relevant for Bayesian calculations of the optimum course of action. Instead of thinking of neural representations as transformations of stimulus energies, we should regard them as approximate estimates of the probable truths of hypotheses about the current environment, for these are the quantities required by a probabilistic brain working on Bayesian principles.

Artificial Intelligence↗

SIMMAP: stochastic character mapping of discrete traits on phylogenies.

BACKGROUND: Character mapping on phylogenies has played an important, if not critical role, in our understanding of molecular, morphological, and behavioral evolution. Until very recently we have relied on parsimony to infer character changes. Parsimony has a number of serious limitations that are drawbacks to our understanding. Recent statistical methods have been developed that free us from these limitations enabling us to overcome the problems of parsimony by accommodating uncertainty in evolutionary time, ancestral states, and the phylogeny. RESULTS: SIMMAP has been developed to implement stochastic character mapping that is useful to both molecular evolutionists, systematists, and bioinformaticians. Researchers can address questions about positive selection, patterns of amino acid substitution, character association, and patterns of morphological evolution. CONCLUSION: Stochastic character mapping, as implemented in the SIMMAP software, enables users to address questions that require mapping characters onto phylogenies using a probabilistic approach that does not rely on parsimony. Analyses can be performed using a fully Bayesian approach that is not reliant on considering a single topology, set of substitution model parameters, or reconstruction of ancestral states. Uncertainty in these quantities is accommodated by using MCMC samples from their respective posterior distributions.

Animals↗

Comments about Joint Modeling of Cluster Size and Binary and Continuous Subunit-Specific Outcomes.

In longitudinal studies and in clustered situations often binary and continuous response variables are observed and need to be modeled together. In a recent publication Dunson, Chen, and Harry (2003, Biometrics 59, 521-530) (DCH) propose a Bayesian approach for joint modeling of cluster size and binary and continuous subunit-specific outcomes and illustrate this approach with a developmental toxicity data example. In this note we demonstrate how standard software (PROC NLMIXED in SAS) can be used to obtain maximum likelihood estimates in an alternative parameterization of the model with a single cluster-level factor considered by DCH for that example. We also suggest that a more general model with additional cluster-level random effects provides a better fit to the data set. An apparent discrepancy between the estimates obtained by DCH and the estimates obtained earlier by Catalano and Ryan (1992, Journal of the American Statistical Association 87, 651-658) is also resolved. The issue of bias in inferences concerning the dose effect when cluster size is ignored is discussed. The maximum-likelihood approach considered herein is applicable to general situations with multiple clustered or longitudinally measured outcomes of different type and does not require prior specification and extensive programming.

Animals↗

Cenozoic biogeography and evolution in direct-developing frogs of Central America (Leptodactylidae: Eleutherodactylus) as inferred from a phylogenetic analysis of nuclear and mitochondrial genes.

We report the first phylogenetic analysis of DNA sequence data for the Central American component of the genus Eleutherodactylus (Anura: Leptodactylidae: Eleutherodactylinae), one of the most ubiquitous, diverse, and abundant components of the Neotropical amphibian fauna. We obtained DNA sequence data from 55 specimens representing 45 species. Sampling was focused on Central America, but also included Bolivia, Brazil, Jamaica, and the USA. We sequenced 1460 contiguous base pairs (bp) of the mitochondrial genome containing ND2 and five neighboring tRNA genes, plus 1300 bp of the c-myc nuclear gene. The resulting phylogenetic inferences were broadly concordant between data sets and among analytical methods. The subgenus Craugastor is monophyletic and its initial radiation was potentially rapid and adaptive. Within Craugastor, the earliest splits separate three northern Central American species groups, milesi, augusti, and alfredi, from a clade comprising the rest of Craugastor. Within the latter clade, the rhodopis group as formerly recognized comprises three deeply divergent clades that do not form a monophyletic group; we therefore restrict the content of the rhodopis group to one of two northern clades, and use new names for the other northern (mexicanus group) and one southern clade (bransfordii group). The new rhodopis and bransfordii groups together form the sister taxon to a clade comprising the biporcatus, fitzingeri, mexicanus, and rugulosus groups. We used a Bayesian MCMC approach together with geological and biogeographic assumptions to estimate divergence times from the combined DNA sequence data. Our results corroborated three independent dispersal events for the origins of Central American Eleutherodactylus: (1) an ancestor of Craugastor entered northern Central America from South American in the early Paleocene, (2) an ancestor of the subgenus Syrrhophus entered northern Central America from the Caribbean at the end of the Eocene, and (3) a wave of independent dispersal events from South America coincided with formation of the Isthmus of Panama during the Pliocene. We elevate the subgenus Craugastor to the genus rank.

Animals↗

Phylogeography of the rock partridge (Alectoris graeca).

We used mitochondrial DNA control-region and microsatellite data to infer the evolutionary history and past demographic changes in 332 rock partridges (Alectoris graeca) sampled from throughout the species' distribution range, with the exception of the central Balkans region. Maternal and biparental DNA markers indicated concordantly that rock partridge populations are structured geographically (mtDNA phiST = 0.86, microsatellite FST = 0.35; RST = 0.31; P < 0.001). Phylogenetic analyses of 22 mtDNA haplotypes identified two major phylogroups (supported by bootstrap values = 93%), splitting partridges from Sicily vs. all the other sampled populations at an average Tamura-Nei genetic distance of 0.035, which corresponds to 65% of the average distance between closely related species of Alectoris. Coalescent estimates of divergence times suggested that rock partridges in Sicily were isolated for more than 200000 years. This deep subdivision was confirmed by multivariate, Bayesian clustering and population assignment analyses of microsatellite genotypes, which supported also a subdivision of partridges from the Alps vs. populations in the Apennines, Albania and Greece. Partridges in the Apennines and Albania-Greece were probably connected by gene flow since recently through a late Pleistocene Adriatic landbridge. Deglaciated Alps were probably colonized by distinct and, perhaps, not yet sampled source populations. Bottleneck and mismatch analyses indicate that rock partridges have lost variability through past population declines, and did not expand recently. Deglaciated areas could have been recolonized without any strong demographic expansion. Genetic data partially supported subspecies subdivisions, and allowed delimiting distinct conservation units. Rock partridges in Sicily, formally recognized as A. g. whitakeri, met the criteria for a distinct evolutionary significant unit.

Animals↗

Conservation genetics and demographic history of the endangered Cape Fear shiner (Notropis mekistocholas).

We examined allelic variation at 22 nuclear-encoded markers (21 microsatellites and one anonymous locus) and mitochondrial (mt)DNA in two geographical samples of the endangered cyprinid fish Notropis mekistocholas (Cape Fear shiner). Genetic diversity was relatively high in comparison to other endangered vertebrates, and there was no evidence of small population effects despite the low abundance reported for the species. Significant heterogeneity (following Bonferroni correction) in allele distribution at three microsatellites and in haplotype distribution in mtDNA was detected between the two localities. This heterogeneity may be due to reduced gene flow caused by a dam built in the early 1900 s. Bayesian coalescent analysis of microsatellite variation indicated that effective population size of Cape Fear shiners has declined in recent times (11-25 435 years ago, with highest posterior probabilities between 126 and 2007 years ago) by one-two orders of magnitude, consistent with the observed decline in abundance of the species. A decline in effective size was not indicated by analysis of mtDNA, where sequence polymorphism appeared to carry the signature of an older expansion phase that dated to the Pleistocene ( approximately 12 700 > 1 million years ago). Cape Fear shiners thus appear to have undergone an expansion phase following a glacial cycle but to have declined significantly in more recent times. These results suggest that rapidly evolving markers such as microsatellites may constitute a suitable tool when inferring recent demographic dynamics of populations.

Animals↗

Inferring parameters of mutation, selection and demography from patterns of synonymous site evolution in Drosophila.

Selection acting on codon usage can cause patterns of synonymous evolution to deviate considerably from those expected under neutrality. To investigate the quantitative relationship between parameters of mutation, selection, and demography, and patterns of synonymous site divergence, we have developed a novel combination of population genetic models and likelihood methods of phylogenetic sequence analysis. Comparing 50 orthologous gene pairs from Drosophila melanogaster and D. virilis and 27 from D. melanogaster and D. simulans, we show considerable variation between amino acids and genes in the strength of selection acting on codon usage and find evidence for both long-term and short-term changes in the strength of selection between species. Remarkably, D. melanogaster shows no evidence of current selection on codon usage, while its sister species D. simulans experiences only half the selection pressure for codon usage of their common ancestor. We also find evidence for considerable base asymmetries in the rate of mutation, such that the average synonymous mutation rate is 20-30% higher than in noncoding regions. A Bayesian approach is adopted to investigate how accounting for selection on codon usage influences estimates of the parameters of mutation.

Animals↗