Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Footprints in the sand: independent reduction of subdigital lamellae in the Namib-Kalahari burrowing geckos.

Many desert organisms exhibit convergence, and certain physical factors such as windblown sands have generated remarkably similar ecomorphs across divergent lineages. The burrowing geckos Colopus, Chondrodactylus and Palmatogecko occupy dune ecosystems in the Namib and Kalahari deserts of southwest Africa. Considered closely related, they share several putative synapomorphies, including reduced subdigital pads (toe pads) and spinose digital scales. Though recognized as part of Africa's ecologically diverse Pachydactylus Group, the burrowing geckos' precise phylogenetic affinities remain elusive. Convergent pedal modification provides a tenable alternative explaining the geckos' derived terrestriality and adaptation to Namib and Kalahari sands. We generated a molecular phylogeny for the Pachydactylus Group to examine evolutionary relationships among the burrowing geckos and infer historical patterns of pedal character change. Bayesian and parsimony analyses revealed all three burrowing genera to be deeply nested within Pachydactylus, each genus belonging to a separate clade. Strong support for these distinct clades indicates ecomorphological adaptations for burrowing have evolved independently three times in the southern Pachydactylus Group. We argue that the physical properties of Namib and Kalahari sands played a principal role in selecting for pedal similarity.

Animals↗

Bayesian isotonic regression and trend analysis.

In many applications, the mean of a response variable can be assumed to be a nondecreasing function of a continuous predictor, controlling for covariates. In such cases, interest often focuses on estimating the regression function, while also assessing evidence of an association. This article proposes a new framework for Bayesian isotonic regression and order-restricted inference. Approximating the regression function with a high-dimensional piecewise linear model, the nondecreasing constraint is incorporated through a prior distribution for the slopes consisting of a product mixture of point masses (accounting for flat regions) and truncated normal densities. To borrow information across the intervals and smooth the curve, the prior is formulated as a latent autoregressive normal process. This structure facilitates efficient posterior computation, since the full conditional distributions of the parameters have simple conjugate forms. Point and interval estimates of the regression function and posterior probabilities of an association for different regions of the predictor can be estimated from a single MCMC run. Generalizations to categorical outcomes and multiple predictors are described, and the approach is applied to an epidemiology application.

Bayes Theorem↗

Population structure of loggerhead shrikes in the California Channel Islands.

The loggerhead shrike (Lanius ludovicianus), a songbird that hunts like a small raptor, maintains breeding populations on seven of the eight California Channel Islands. One of the two subspecies, L. l. anthonyi, was described as having breeding populations on six of the islands while a second subspecies, L. l. mearnsi, was described as being endemic to San Clemente Island. Previous genetic studies have demonstrated that the San Clemente Island loggerhead shrike is well differentiated genetically from both L. l. anthonyi and mainland populations, despite the fact that birds from outside the population are regular visitors to the island. Those studies, however, did not include a comparison between San Clemente Island shrikes and the breeding population on Santa Catalina Island, the closest island to San Clemente. Here we use mitochondrial control region sequences and nuclear microsatellites to investigate the population structure of loggerhead shrikes in the Channel Islands. We confirm the genetic distinctiveness of the San Clemente Island loggerhead shrike and, using Bayesian clustering analysis, demonstrate the presence and infer the source of the nonbreeding visitors. Our results indicate that Channel Island loggerhead shrikes comprise three distinct genetic clusters that inhabit: (i) San Clemente Island, (ii) Santa Catalina Island and (iii) the Northern Channel Islands and nearby mainland; they do not support a recent suggestion that all Channel Island loggerhead shrikes should be managed as a single entity.

Animals↗

Range expansions in the flightless longhorn cactus beetles, Moneilema gigas and Moneilema armatum, in response to Pleistocene climate changes.

Pollen cores and plant and animal fossils suggest that global climate changes at the end of the last glacial period caused range expansions in organisms indigenous to the North American desert regions, but this suggestion has rarely been investigated from a population genetic perspective. In order to investigate the impact of Pleistocene climate changes and glacial/interglacial cycling on the distribution and population structure of animals in North American desert communities, biogeographical patterns in the flightless, warm-desert cactus beetles, Moneilema gigas and Moneilema armatum, were examined using mitochondrial DNA (mtDNA) sequence data from the cytochrome oxidase I (COI) gene. Gene tree relationships between haplotypes were inferred using parsimony, maximum-likelihood, and Bayesian analysis. Nested clade analysis and coalescent modelling using the programs mdiv and fluctuate were used to identify demographically independent populations, and to test the hypothesis that Pleistocene climate changes caused recent range expansions in these species. A sign test was used to evaluate the probability of observing concerted population growth across multiple, independent populations. The phylogeographical and nested clade analyses reveal a history of northward expansion in both of these species, as well as a history of past range fragmentation, followed by expansion from refugia. The coalescent analyses provide highly significant evidence for independent range expansions from multiple refugia, but also identify biogeographical patterns that predate the most recent glacial period. The results indicate that widespread desert environments are more ancient than has been suggested in the past.

Animals↗

Phylogeography of the longhorn cactus beetle Moneilema appressum LeConte (Coleoptera: Cerambycidae): was the differentiation of the Madrean sky islands driven by Pleistocene climate changes?

Although it has been suggested that Pleistocene climate changes drove population differentiation and speciation in many groups of organisms, population genetic evidence in support of this scenario has been ambiguous, and it has often been difficult to distinguish putative vicariance from simple isolation by distance. The sky island communities of the American Southwest present an ideal system in which to compare late Pleistocene range fragmentations documented by palaeoenvironmental studies with population genetic data from organisms within these communities. In order to elucidate the impact of Pleistocene climate fluctuations on these environments, biogeographic patterns in the flightless longhorn cactus beetle, Moneilema appressum were examined using mitochondrial DNA sequence data. Gene tree relationships between haplotypes were inferred using parsimony, maximum-likelihood, and Bayesian analysis. Nested clade analysis, Mantel tests, and coalescent modelling were employed to examine alternative biogeographic scenarios, and to test the hypothesis that Pleistocene climate changes drove population differentiation in this species. The program mdiv was used to estimate migration and divergence times between populations, and to measure the statistical support for isolation over ongoing migration. These analyses showed significant geographic structure in genetic relationships, and implicated topography as a key determinant of isolation. However, although the coalescent analyses suggested that a history of past habitat fragmentation underlies the observed geographic patterns, the nested clade analysis indicated that the pattern was consistent with isolation by distance. Estimated divergence times indicated that range fragmentation in M. appressum is considerably older than the end of the most recent glacial, but coincided with earlier interglacial warming events and with documented range expansions in other, desert-dwelling species of Moneilema.

Animals↗

Modeling population kinetics of free fatty acids in isolated rat hepatocytes using Markov Chain Monte Carlo.

The aim of this study is the characterization, by means of mathematical models, of the activity of isolated hepatic rat cells as regards the conversion of free fatty acids (FFA) to ketone bodies (KB). A new physiologically based compartmental model of FFA metabolism is used within a context of population pharmacokinetics. This analysis is based on a hierarchical model, that differs from standard model formulations, to account for the fact that some data sets belong to the same animal but have been collected under different experimental conditions. The statistical inference problem has been addressed within a Bayesian context and solved by using Markov Chain Monte Carlo (MCMC) simulation. The results obtained in this study indicate that, although hormones epinephrine and insulin are important metabolic regulatory factors in vivo, the conversion of FFA to KB by isolated hepatic rat cells is not significantly affected by epinephrine and only little influenced by insulin. So we conclude that in vivo, the interaction of these two hormones with other compounds not considered in this study plays a fundamental role in ketogenesis. From this study it appears that mathematical models of metabolic processes can be successfully employed in population kinetic studies using MCMC methods.

Animals↗

Genomic exploration of the journey of Plasmodium vivax in Latin America.

Plasmodium vivax is the predominant malaria parasite in Latin America. Its colonization history in the region is rich and complex, and is still highly debated, especially about its origin(s). Our study employed cutting-edge population genomic techniques to analyze whole genome variation from 620 P. vivax isolates, including 107 newly sequenced samples from West Africa, Middle East, and Latin America. This sampling represents nearly all potential source populations worldwide currently available. Analyses of the genetic structure, diversity, ancestry, coalescent-based inferences, including demographic scenario testing using Approximate Bayesian Computation, have revealed a more complex evolutionary history than previously envisioned. Indeed, our analyses suggest that the current American P. vivax populations predominantly stemmed from a now-extinct European lineage, with the potential contribution also from unsampled populations, most likely of West African origin. We also found evidence that P. vivax arrived in Latin America in multiple waves, initially during early European contact and later through post-colonial human migration waves in the late 19th-century. This study provides a fresh perspective on P. vivax's intricate evolutionary journey and brings insights into the possible contribution of West African P. vivax populations to the colonization history of Latin America.

Plasmodium vivax↗

A quantitative genetic method for estimating developmental instability.

The concept of developmental instability (DI) is frequently used in evolutionary biology, and a range of definitions has been proposed. Moreover, numerous different statistical methods have been used for estimation of DI. The common basis for all methods is that measures need to be obtained from repeated structures within organisms. In the case of fluctuating asymmetry, mirror images could be interpreted as the repeats of each other. All repeats of a trait on one organism should, from a quantitative perspective, have the same genetic foundation. Most previous methods have not accounted for the genetics of the underlying trait. It is here shown how a statistical method from quantitative genetics (the repeated records animal model) can be used for assessment of DI, based on estimation of the variance due to the permanent environment. Moreover, Gibbs sampling is used for inference of the parameters, which provides a Bayesian framework where posterior distributions easily can be calculated from any functions of the variance components. The method is applied to a real dataset from two populations of the plant Scabiosa canescens, and results shows that it works well under realistic situations.

Analysis of Variance↗

Detecting the historical signature of key innovations using stochastic models of character evolution and cladogenesis.

Phylogenetic evidence for biological traits that increase the net diversification rate of lineages (key innovations) is most commonly drawn from comparisons of clade size. This can work well for ancient, unreversed traits and for correlating multiple trait origins with higher diversification rates, but it is less suitable for unique events, recently evolved innovations, and traits that exhibit homoplasy. Here I present a new method for detecting the phylogenetic signature of key innovations that tests whether the evolutionary history of the candidate trait is associated with shorter waiting times between cladogenesis events. The method employs stochastic models of character evolution and cladogenesis and integrates well into a Bayesian framework in which uncertainty in historical inferences (such as phylogenetic relationships) is allowed. Applied to a well-known example in plants, nectar spurs in columbines, the method gives much stronger support to the key innovation hypothesis than previous tests.

Adaptation, Biological↗

Bayesian mapping of quantitative trait loci under the identity-by-descent-based variance component model.

Variance component analysis of quantitative trait loci (QTL) is an important strategy of genetic mapping for complex traits in humans. The method is robust because it can handle an arbitrary number of alleles with arbitrary modes of gene actions. The variance component method is usually implemented using the proportion of alleles with identity-by-descent (IBD) shared by relatives. As a result, information about marker linkage phases in the parents is not required. The method has been studied extensively under either the maximum-likelihood framework or the sib-pair regression paradigm. However, virtually all investigations are limited to normally distributed traits under a single QTL model. In this study, we develop a Bayes method to map multiple QTL. We also extend the Bayesian mapping procedure to identify QTL responsible for the variation of complex binary diseases in humans under a threshold model. The method can also treat the number of QTL as a parameter and infer its posterior distribution. We use the reversible jump Markov chain Monte Carlo method to infer the posterior distributions of parameters of interest. The Bayesian mapping procedure ends with an estimation of the joint posterior distribution of the number of QTL and the locations and variances of the identified QTL. Utilities of the method are demonstrated using a simulated population consisting of multiple full-sib families.

Alleles↗

Phylogenetic reconstruction of a known HIV-1 CRF04_cpx transmission network using maximum likelihood and Bayesian methods.

The CRF04_cpx strains of HIV-1 accounts for approximately 2-10% of the infected population in Greece, across different transmission risk groups. CRF04_cpx was the lineage documented in an HIV-1 transmission network in Thessalonica, northern Greece. Most of the transmissions occurred through unprotected heterosexual contacts between 1989 and 1993. Blood samples were available for six patients, obtained 6-10 years later, except for one patient sampled in 1991. Our objective was to examine whether the transmission history is compatible with the evolutionary tree of the virus, in partial gag, partial env, and partial gag+env. The inferred phylogenetic tree obtained using maximum likelihood and Bayesian methods in partial gag+env was much closer to the transmission tree than that using either env or gag separately. Our findings suggest that the epidemiological relationships among patients who have been infected by a common source correspond almost exactly to the evolutionary trees of the virus, given that enough phylogenetic signal is present in the alignment. Moreover, we found evidence that recombination is not the most parsimonious explanation for the phylogenetic incongruence between gag and env. For patients with known infection dates, the estimated dates of the coalescent events obtained using molecular clock calculations based on a newly developed Bayesian method in gag + env were in agreement with the actual infection dates.

Bayes Theorem↗

Statistical evaluation of pairwise protein sequence comparison with the Bayesian bootstrap.

MOTIVATION: Protein sequence comparison methods are routinely used to infer the intricate network of evolutionary relationships found within the rapidly growing library of protein sequences, and thereby to predict the structure and function of uncharacterized proteins. In the present study, we detail an improved statistical benchmark of pairwise protein sequence comparison algorithms. We use bootstrap resampling techniques to determine standard statistical errors and to estimate the confidence of our conclusions. We show that the underlying structure within benchmark databases causes Efron's standard, non-parametric bootstrap to be biased. Consequently, the standard bootstrap underpredicts average performance when used in the context of evaluating sequence comparison methods. We have developed, as an alternative, an unbiased statistical evaluation based on the Bayesian bootstrap, a resampling method operationally similar to the standard bootstrap. RESULTS: We apply our analysis to the comparative study of amino acid substitution matrix families and find that using modern matrices results in a small, but statistically significant improvement in remote homology detection compared with the classic PAM and BLOSUM matrices. AVAILABILITY: The sequence sets and code for performing these analyses are available from http://compbio.berkeley.edu/. CONTACT: brenner@compbio.berkeley.edu.

Algorithms↗

Sparse polygenic risk score inference with the spike-and-slab LASSO.

MOTIVATION: Large-scale biobanks, with rich phenotypic and genomic data across hundreds of thousands of samples, provide ample opportunities to elucidate the genetics of complex traits and diseases. Consequently, there is growing demand for robust and scalable methods for disease risk prediction from genotype data. Inference in this setting is challenging due to the high-dimensionality of genomic data, especially when coupled with smaller sample sizes. Popular Polygenic Risk Score (PRS) inference methods address this challenge by adopting sparse Bayesian priors or penalized regression techniques, such as the Least Absolute Shrinkage and Selection Operator (LASSO). However, the former class of methods are not as scalable and do not produce exact sparsity, while the latter tends to over-shrink large coefficients. RESULTS: In this study, we present SSLPRS, a novel PRS method based on the Spike-and-Slab LASSO (SSL) prior, which offers a theoretical bridge between the two frameworks. We extend previous work to derive a coordinate-ascent inference algorithm that operates on GWAS summary statistics, which is orders-of-magnitude more efficient than corresponding individual-level-based implementations. To illustrate the statistical properties of the proposed model, we conducted experiments involving nine simulation configurations and nine quantitative phenotypes from the UK Biobank. Our results demonstrate that SSLPRS is competitive with state-of-the-art methods in terms of prediction accuracy and exhibits superior variable selection performance, especially in sparse genetic architectures. In simulations, this translates to upwards of 50% improvement in positive predictive value. In analysis of real phenotypes, we show that selected variants are highly enriched for meaningful genomic annotations and have better replication rates in larger meta-analyses. AVAILABILITY AND IMPLEMENTATION: SSLPRS is available in the open-source package https://github.com/li-lab-mcgill/penprs.

Multifactorial Inheritance↗

An epidemiologic critique of current microbial risk assessment practices: the importance of prevalence and test accuracy data.

Data deficiencies are impeding the development and validation of microbial risk assessment models. One such deficiency is the failure to adjust test-based (apparent) prevalence estimates to true prevalence estimates by correcting for the imperfect accuracy of tests that are used. Such adjustments will facilitate comparability of data from different populations and from the same population over time as tests change and the unbiased quantification of effects of mitigation strategies. True prevalence can be estimated from apparent prevalence using frequentist and Bayesian methods, but the latter are more flexible and can incorporate uncertainty in test accuracy and prior prevalence data. Both approaches can be used for single or multiple populations, but the Bayesian approach can better deal with clustered data, inferences for rare events, and uncertainty in multiple variables. Examples of prevalence inferences based on results of Salmonella culture are presented. The opportunity to adjust test-based prevalence estimates is predicated on the availability of sensitivity and specificity estimates. These estimates can be obtained from studies using archived gold standard (reference) samples, by screening with the new test and follow-up of test-positive and test-negative samples with a gold standard test, and by use of latent class methods, which make no assumptions about the true status of each sampling unit. Latent class analysis can be done with maximum likelihood and Bayesian methods, and an example of their use in the evaluation of tests for Toxoplasma gondii in pigs is presented. Guidelines are proposed for more transparent incorporation of test data into microbial risk assessments.

Bayes Theorem↗

Constructing probabilistic models.

Bayesian networks have become one of the most popular probabilistic techniques in AI, largely due to the development of several efficient inference algorithms. In this paper we describe a heuristic method for constructing Bayesian networks. Our construction method relies on the relationship between Bayesian networks and decomposable models, a special kind of graphical model. We explain this relationship and then show how it can be used to facilitate model construction. Finally, we describe an implemented computer program that illustrates these ideas.

Algorithms↗

Empirical and hierarchical Bayesian estimation of ancestral states.

Several methods have been proposed to infer the states at the ancestral nodes on a phylogeny. These methods assume a specific tree and set of branch lengths when estimating the ancestral character state. Inferences of the ancestral states, then, are conditioned on the tree and branch lengths being true. We develop a hierarchical Bayes method for inferring the ancestral states on a tree. The method integrates over uncertainty in the tree, branch lengths, and substitution model parameters by using Markov chain Monte Carlo. We compare the hierarchical Bayes inferences of ancestral states with inferences of ancestral states made under the assumption that a specific tree is correct. We find that the methods are correlated, but that accommodating uncertainty in parameters of the phylogenetic model can make inferences of ancestral states even more uncertain than they would be in an empirical Bayes analysis.

Animals↗

Inferring gene regulatory networks from time series data using the minimum description length principle.

MOTIVATION: A central question in reverse engineering of genetic networks consists in determining the dependencies and regulating relationships among genes. This paper addresses the problem of inferring genetic regulatory networks from time-series gene-expression profiles. By adopting a probabilistic modeling framework compatible with the family of models represented by dynamic Bayesian networks and probabilistic Boolean networks, this paper proposes a network inference algorithm to recover not only the direct gene connectivity but also the regulating orientations. RESULTS: Based on the minimum description length principle, a novel network inference algorithm is proposed that greatly shrinks the search space for graphical solutions and achieves a good trade-off between modeling complexity and data fitting. Simulation results show that the algorithm achieves good performance in the case of synthetic networks. Compared with existing state-of-the-art results in the literature, the proposed algorithm exceptionally excels in efficiency, accuracy, robustness and scalability. Given a time-series dataset for Drosophila melanogaster, the paper proposes a genetic regulatory network involved in Drosophila's muscle development. AVAILABILITY: Available from the authors upon request.

Algorithms↗

[Stereotype-based expectancy and social judgment: rethinking from a Bayesian perspective].

To examine effect of prior stereotypical expectancy on social judgment from a Bayesian perspective, undergraduate subjects (N = 204) were asked to infer a target person's attitude toward an atomic power problem. Half of them were told in advance that he was a member of Liberal Democratic Party (pro-expectancy condition), and the other half were told that he was a member of Japanese Socialist Party (con-expectancy condition). Then subjects were given a series of his previous relevant utterances, which had either high or low diagnostic values for the inference of his attitude. "Labeling effect" occurred. That is, despite being given identical utterances, subjects given L.D.P. label estimated the target's attitude to be more favorable toward the atomic power than subjects given J.S.P. label. This effect emerged mainly when subjects were given low diagnostic utterances. Subjects given high diagnostic utterances inadequately underused the base-rate information (prior expectancy) compared with the Bayesian normative value. Utterances congruent with prior expectancy were better recalled than utterances incongruent with prior expectancy.

Adult↗