Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

The effects of rate variation on ancestral inference in the coalescent.

We describe a Markov chain Monte Carlo approach for assessing the role of site-to-site rate variation in the analysis of within-population samples of DNA sequences using the coalescent. Our framework is a Bayesian one. We discuss methods for assessing the goodness-of-fit of these models, as well as problems concerning the separate estimation of effective population size and mutation rate. Using a mitochondrial data set for illustration, we show that ancestral inference concerning coalescence times can be dramatically affected if rate variation is ignored.

Algorithms↗

Molecular phylogeny and evolution of the freshwater eels genus Anguilla based on the whole mitochondrial genome sequences.

Molecular phylogenetic analyses were conducted using the whole mitochondrial genome sequences of all 18 species/subspecies of the freshwater eels genus Anguilla to infer their phylogenetic relationships and to evaluate hypotheses about the possible dispersal routes of this genus. The Bayesian and maximum likelihood analyses using a total of 15,187 sites of mitochondrial DNA sequences suggested that A. mossambica was the most basal species of anguillid eel, and that the other species (except for A. borneensis) formed three geographic clades: Atlantic (two species), Oceania (three species), and Indo-Pacific (11 species). The present study clearly indicated a sister relationship between the Atlantic and Oceanian species, which now have distantly separated geographic distributions. This suggests that the previous hypotheses to estimate the dispersal route of anguillid eels into the Atlantic Ocean based on the current geographic distribution of species are unsupported by the present more complete analysis. Alternatively, the unique geographic distribution of the present day species in the genus Anguilla appears to have resulted from multiple dispersal events. Although the age of the beginning of speciation among anguillid eels was tentatively estimated as 20 million years ago using a calibration for bony fishes of 7.3x10(-4) substitutions/site/million years, it is possible that this divergence time was underestimated because of the ecological characteristics of these fishes. The results of the present study suggest that the hypotheses for the dispersal route and divergence time of the genus Anguilla should be reconsidered.

Anguilla↗

New DNA data from a transthyretin nuclear intron suggest an Oligocene to Miocene diversification of living South America opossums (Marsupialia: Didelphidae).

Phylogenetic relationships of 19 species of didelphid marsupials were studied using two nuclear markers, the non-coding transthyretin intron 1 (TTR) and the coding interphotoreceptor retinoid binding protein exon 1 (IRBP), and two mitochondrial genes, the protein-coding cytochrome b (cyt-b) and the structural 12S ribosomal DNA (12S rDNA). Evolutionary dynamics of these four markers were compared to each other, revealing the appropriate properties presented by TTR intron 1 together with its well supported and resolved phylogenetic signal. Nuclear markers supported the monophyly of medium and large-sized opossums Metachirus+(Chironectes, Lutreolina, Didelphis, Philander), and the paraphyly of mouse-sized opossums, with the genera Gracilinanus, Thylamys, and Marmosops as a sister group to medium and large-sized didelphids. Conflicting branching patterns between mitochondrial and nuclear data involved the phylogenetic position of Marmosa-Micoureus-Monodelphis relative to other mouse-sized opossums. Nuclear phylogenetic inferences among genera were confirmed by the presence of synapomorphic indels observed in TTR intron 1. A Bayesian relaxed molecular clock dating of didelphid evolution using nuclear markers estimated their origin in the Middle Eocene (39.8 million years ago), with subsequent diversification during the Oligocene (Deseadan) and Miocene.

Animals↗

Multivariate autoregressive modeling of fMRI time series.

We propose the use of multivariate autoregressive (MAR) models of functional magnetic resonance imaging time series to make inferences about functional integration within the human brain. The method is demonstrated with synthetic and real data showing how such models are able to characterize interregional dependence. We extend linear MAR models to accommodate nonlinear interactions to model top-down modulatory processes with bilinear terms. MAR models are time series models and thereby model temporal order within measured brain activity. A further benefit of the MAR approach is that connectivity maps may contain loops, yet exact inference can proceed within a linear framework. Model order selection and parameter estimation are implemented by using Bayesian methods.

Algorithms↗

Accurate gene diversity estimates from amplified fragment length polymorphism (AFLP) markers.

Three procedures for the estimation of null allele frequencies and gene diversity from dominant multilocus data were empirically tested in natural populations of the outcrossing angiosperm Persoonia mollis (Proteaceae). The three procedures were the square root transform of the null homozygote frequency, the Lynch & Milligan procedure, and the Bayesian method. Genotypes for each of 116 polymorphic loci generated by amplified fragment length polymorphism (AFLP) were inferred from segregation patterns in progeny arrays. Therefore, for the plus phenotype (band present), heterozygotes were distinguished from homozygotes. In contrast to previous studies, all three procedures produced very similar mean estimates of heterozygosity, which were in turn accurate estimators of the direct value (HO = 0.28). A second population of P. mollis displayed markedly lower levels of heterozygosity (HO = 0.20) but approximately twice as many polymorphic loci (284). These AFLP results show that biases in estimates of average null allele frequency and heterozygosity are largely eliminated in highly polymorphic dominant marker data sets displaying a J-shaped beta distribution with a high percentage of loci containing more than three null homozygotes and relatively few loci with no null homozygotes. This distribution may be typical of outcrossing angiosperms.

Alleles↗

Bayesian extrapolation of space-time trends in cancer registry data.

We apply a full Bayesian model framework to a dataset on stomach cancer mortality in West Germany. The data are stratified by age group, year, and district. Using an age-period-cohort model with an additional spatial component, our goal is to investigate whether there is evidence for space-time interactions in these data. Furthermore, we will determine whether a period-space or a cohort-space interaction model is more appropriate to predict future mortality rates. The setup will be fully Bayesian based on a series of Gaussian Markov random field priors for each of the components. Statistical inference is based on efficient algorithms to block update Gaussian Markov random fields, which have recently been proposed in the literature.

Bayes Theorem↗

When wrong predictions provide more support than right ones.

Correct predictions of rare events are normatively more supportive of a theory or hypothesis than correct predictions of common ones. In other words, correct bold predictions provide more support than do correct timid predictions. Are lay hypothesis testers sensitive to the boldness of predictions? Results reported here show that participants were very sensitive to boldness, often finding incorrect bold predictions more supportive than correct timid ones. Participants were willing to tolerate inaccurate predictions only when predictions were bold. This finding was demonstrated in the context of competing forecasters and in the context of competing scientific theories. The results support recent views of human inference that postulate that lay hypothesis testers are sensitive to the rarity of data. Furthermore, a normative (Bayesian) account can explain the present results and provides an alternative interpretation of similar results that have been explained using a purely descriptive model.

Adult↗

An evaluation of explanations of probabilistic inference.

Providing explanations of the conclusions of decision-support systems can be viewed as presenting inference results in a manner that enhances the user's insight into how these results were obtained. The ability to explain inferences has been demonstrated to be an important factor in making medical decision-support systems acceptable for clinical use. Although many researchers in artificial intelligence have explored the automatic generation of explanations for decision-support systems based on symbolic reasoning, research in automated explanation of probabilistic results has been limited. We present the results of an evaluation study of INSITE, a program that explains the reasoning of decision-support systems based on Bayesian belief networks. In the domain of anesthesia, we compared subjects who had access to a belief network with explanations of the inference results to control subjects who used the same belief network without explanations. We show that, compared to control subjects, the explanation subjects demonstrated greater diagnostic accuracy, were more confident about their conclusions, were more critical of the belief network, and found the presentation of the inference results more clear.

Bayes Theorem↗

An evaluation of explanations of probabilistic inference.

Providing explanations of the conclusions of decision-support systems can be viewed as presenting inference results in a manner that enhances the user's insight into how these results were obtained. The ability to explain inferences has been demonstrated to be an important factor in making medical decision-support systems acceptable for clinical use. Although many researchers in artificial intelligence have explored the automatic generation of explanations for decision-support systems based on symbolic reasoning, research in automated explanation of probabilistic results has been limited. We present the results of an an evaluation study of INSITE, a program that explains the reasoning of decision-support systems based on Bayesian belief networks. In the domain of anesthesia, we compared subjects who had access to a belief network with explanations of the inference results, to control subjects who used the same belief network without explanations. We show that, compared to control subjects, the explanation subjects demonstrated greater diagnostic accuracy, were more confident about their conclusions, were more critical of the belief network, and found the presentation of the inference results more clear.

Anesthesia↗

Multiscale and Bayesian approaches to data analysis in genomics high-throughput screening.

Tremendous amounts of data are produced by high-throughput screening methods currently employed in drug discovery and product development. A typical cDNA microarray or oligonucleotide-based gene chip experiment easily generates over 10,000 data points for each array or chip. The challenge of inferring meaningful information is formidable given the size and number of these datasets. This paper reviews the current status of statistical tools available for gene expression analysis, with emphasis on Bayesian approaches and multiscale wavelet filtering. Fundamental concepts of Bayesian and multiscale modeling are discussed from the perspective of their potential to address important issues related to the analysis of gene expression data, such as the fact that genomic data often have non-Gaussian distributions and feature localization and multiple scales in both frequency and measurement dimension. Recent publications in these areas are reviewed. Wavelet filtering and the advantages of multiscale methods are demonstrated by application to publicly available gene expression data from the National Cancer Institute (NCI). Multiscale methods, including multiscale principal component analysis (MSPCA), are applied to extract gene subsets and to visualize data in multidimensions for comparisons. Similarity in cell lines and gene selection are effectively visualized and quantitatively compared.

Animals↗

Bayesian communication: a clinically significant paradigm for electronic publication.

OBJECTIVE: To develop a model for Bayesian communication to enable readers to make reported data more relevant by including their prior knowledge and values. BACKGROUND: To change their practice, clinicians need good evidence, yet they also need to make new technology applicable to their local knowledge and circumstances. Availability of the Web has the potential for greatly affecting the scientific communication process between research and clinician. Going beyond format changes and hyperlinking, Bayesian communication enables readers to make reported data more relevant by including their prior knowledge and values. This paper addresses the needs and implications for Bayesian communication. FORMULATION: Literature review and development of specifications from readers', authors', publishers', and computers' perspectives consistent with formal requirements for Bayesian reasoning. RESULTS: Seventeen specifications were developed, which included eight for readers (express prior knowledge, view effect size and variability, express threshold, make inferences, view explanation, evaluate study and statistical quality, synthesize multiple studies, and view prior beliefs of the community), three for authors (protect the author's investment, publish enough information, make authoring easy), three for publishers (limit liability, scale up, and establish a business model), and two for computers (incorporate into reading process, use familiar interface metaphors). A sample client-only prototype is available at http://omie.med.jhmi.edu/bayes. CONCLUSION: Bayesian communication has formal justification consistent with the needs of readers and can best be implemented in an online environment. Much research must be done to establish whether the formalism and the reality of readers' needs can meet.

Bayes Theorem↗

Bayesian analysis of quantitative antimicrobial assays.

Quantitative antimicrobial assays are used to assess the efficacy of chemical germicides. Standard methods for statistical analysis use log reduction (LR), the difference on the log scale between average surviving microbes for control and test carriers, as an efficacy measure. These methods have several deficiencies. The LR parameter is not on the original response scale, which complicates its interpretation. The presence of two different definitions of LR makes the statistical inference even more difficult. Current statistical methods for antimicrobial assay analysis rely on asymptotic normal theory, which might not work well for small samples. In addition, they do not appropriately incorporate censored ('too numerous to be counted') observations in the analysis. To overcome those problems, a new Bayesian approach is introduced here. It has also the advantages of more flexible statistical inference, and incorporated prior information in the model.

Bacteria↗

Within-category feature correlations and Bayesian adjustment strategies.

To the extent that categories inform judgments about items, the accuracy with which categories capture the statistical structure of experience should affect judgment accuracy. The authors argue that representations of feature correlations can serve as Bayesian priors, increasing the accuracy of stimulus estimates by decreasing variability. Participants viewed a series of objects that varied on two dimensions that were either uncorrelated or correlated. They estimated each item by manipulating a response object to make it match the presented stimulus. Subsequent classification and feature-inference tasks indicated that the correlation was detected. The pattern of variability in recollections of stimuli suggested that the feature correlation informed estimates as predicted by a Bayesian model of category effects on memory.

Bayes Theorem↗

Bayesian analysis of botanical epidemics using stochastic compartmental models.

A stochastic model for an epidemic, incorporating susceptible, latent, and infectious states, is developed. The model represents primary and secondary infection rates and a time-varying host susceptibility with applications to a wide range of epidemiological systems. A Markov chain Monte Carlo algorithm is presented that allows the model to be fitted to experimental observations within a Bayesian framework. The approach allows the uncertainty in unobserved aspects of the process to be represented in the parameter posterior densities. The methods are applied to experimental observations of damping-off of radish (Raphanus sativus) caused by the fungal pathogen Rhizoctonia solani, in the presence and absence of the antagonistic fungus Trichoderma viride, a biological control agent that has previously been shown to affect the rate of primary infection by using a maximum-likelihood estimate for a simpler model with no allowance for a latent period. Using the Bayesian analysis, we are able to estimate the latent period from population data, even when there is uncertainty in discriminating infectious from latently infected individuals in data collection. We also show that the inference that T. viride can control primary, but not secondary, infection is robust to inclusion of the latent period in the model, although the absolute values of the parameters change. Some refinements and potential difficulties with the Bayesian approach in this context, when prior information on parameters is lacking, are discussed along with broader applications of the methods to a wide range of epidemiological systems.

Algorithms↗

Estimating countermeasure effects for reducing collisions at highway-railway grade crossings.

Frequently transportation engineers are required to make difficult safety investment decisions in the face of uncertainty concerning the cost-effectiveness of different countermeasures. For certain types of highway-railway grade crossings, this problem is further aggravated due to the lack of observed before and after collision data that reflects the impact of specific countermeasures. This study proposes a Bayesian data fusion method as an attempt to overcome these challenges. In this framework, we make use of previous research findings on the effectiveness of a given countermeasure, which could vary by jurisdictions and operating conditions to obtain a priori inference on its expected effects. We then use locally calibrated models, which are valid for a specific jurisdiction, to develop the current best estimates regarding the countermeasure effects. By using a Bayesian framework, these two sources are integrated to obtain the posterior distribution of the countermeasure effectiveness. As a result, the outputs provide information not only of the expected collision response to a specific countermeasure but also its variance and corresponding probability distribution for a range of likely values. Examples from Canadian highway-railway grade crossing data are used to illustrate the proposed methodology and the specific effects of prior knowledge and data likelihood on the combined estimates of countermeasure effects.

Accident Prevention↗

Bayesian analysis of the differences of count data.

Paired count data usually arise in medicine when before and after treatment measurements are considered. In the present paper we assume that the correlated paired count data follow a bivariate Poisson distribution in order to derive the distribution of their difference. The derived distribution is shown to be the same as the one derived for the difference of the independent Poisson variables, thus recasting interest on the distribution introduced by Skellam. Using this distribution we remove correlation, which naturally exists in paired data, and we improve the quality of our inference by using exact distributions instead of normal approximations. The zero-inflated version is considered to account for an excess of zero counts. Bayesian estimation and hypothesis testing for the models considered are discussed. An example from dental epidemiology is used to illustrate the proposed methodology.

Algorithms↗

Inference for smooth curves in longitudinal data with application to an AIDS clinical trial.

We discuss a longitudinal study where data for many subjects are collected at irregular intervals. The study is a randomized trial of HIV infected subjects and the response variable of interest is serum neopterin. The mean of the outcome variable, taken over patients in each treatment group, is assumed to follow a smooth curve. Piecewise cubic polynomials with a moderate number of knots are used to model the curves. A general parametric form is assumed for the covariance structure. Maximum penalized likelihood estimation is used to smooth the over-parameterized curves. Statistical inference for the mean curves, including confidence bands and hypothesis tests, is discussed. Two approaches, one using a Bayesian interpretation of the penalized likelihood and the other based on the asymptotic distribution of the maximum penalized likelihood estimates, are discussed and contrasted. The properties of the confidence bands obtained from these two approaches are evaluated by examining their coverage rates in a simulation study.

Acquired Immunodeficiency Syndrome↗

Population-level history of the wrentit (Chamaea fasciata): implications for comparative phylogeography in the California Floristic Province.

The phylogeography of a variety of species has been studied within the California Floristic Province; however, few studies have examined genetic variation in bird species across the entire region. This study uses mitochondrial DNA data to investigate the phylogeography of the wrentit (Chamaea fasciata), a sedentary bird native to scrub and chaparral habitats of this region. Analysis of molecular variance shows geographic structure, and maximum likelihood, Bayesian, and parsimony analyses consistently identify six main clades that are each restricted geographically. Nested clade phylogeographic analyses infer an overall range expansion for the entire cladogram, and a range expansion is also inferred from the mismatch distribution. Thus, our results suggest that the wrentit was isolated into southern refugia during the Pleistocene and has undergone a recent range expansion. Southern refugia and a range expansion were also identified in a previous study of the California thrasher (Toxostoma redivivum). The wrentit did not show marked divergence between northern and southern California defined by the Transverse Ranges, a pattern seen in a variety of other taxa within this region, including some birds.

Animals↗