Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Hierarchical models in generalized synthesis of evidence: an example based on studies of breast cancer screening.

Evidence regarding the potential benefits of a particular health care intervention is often available from a variety of disparate sources. However, formal synthesis of such evidence has traditionally concentrated almost exclusively on that derived from randomized studies, although for a range of conditions the randomized evidence will be less than adequate due to economic, organizational or ethical considerations. In such situations a formal synthesis of the evidence that is available from observational studies can be valuable whilst awaiting higher quality evidence from randomized trials. Consideration of randomized studies alone may be appropriate when assessing the efficacy of an intervention, but assessment of the effectiveness of such an intervention within a more general target population may be improved by consideration of evidence from non-randomized studies as well. Standard meta-analysis methods may allow for both within- and between-study heterogeneity; however when multiple sources of evidence are considered an extra level of complexity is introduced, namely study type. One possible solution to the problem of making inferences, particularly regarding an overall population effect, in such situations is to model the heterogeneity, both quantitative and qualitative, using a Bayesian hierarchical model. The hierarchical nature of such models specifically allows for the quantitative within and between sources of heterogeneity, whilst the Bayesian approach can accommodate a priori beliefs regarding qualitative differences between the various sources of evidence. The use of such methods in practice is illustrated in the context of screening for breast cancer; in this example evidence is available from both randomized clinical trials and observational studies. A particular appeal of a Bayesian approach for this type of problem lies in the prediction of future benefits likely to be observed in a target population. This approach to health service monitoring in general is discussed.

Aged↗

Combination of direct and indirect evidence in mixed treatment comparisons.

Mixed treatment comparison (MTC) meta-analysis is a generalization of standard pairwise meta-analysis for A vs B trials, to data structures that include, for example, A vs B, B vs C, and A vs C trials. There are two roles for MTC: one is to strengthen inference concerning the relative efficacy of two treatments, by including both 'direct' and 'indirect' comparisons. The other is to facilitate simultaneous inference regarding all treatments, in order for example to select the best treatment. In this paper, we present a range of Bayesian hierarchical models using the Markov chain Monte Carlo software WinBUGS. These are multivariate random effects models that allow for variation in true treatment effects across trials. We consider models where the between-trials variance is homogeneous across treatment comparisons as well as heterogeneous variance models. We also compare models with fixed (unconstrained) baseline study effects with models with random baselines drawn from a common distribution. These models are applied to an illustrative data set and posterior parameter distributions are compared. We discuss model critique and model selection, illustrating the role of Bayesian deviance analysis, and node-based model criticism. The assumptions underlying the MTC models and their parameterization are also discussed.

Humans↗

Further statistics in dentistry. Part 9: Bayesian statistics.

Statistics can be defined as the methods used to assimilate data, so that guidance can be given, and conclusions drawn, in situations which involve uncertainty. In particular, statistical inference is concerned with drawing conclusions about particular aspects of a population when that population cannot be studied in full. Uncertainty arises here because the totality of the information is not available. Instead, to make inferences about the population, it is necessary to rely on a sample of data which is selected from the population; this sample data may be augmented, in certain circumstances, by auxiliary information which is obtained independently of the sample data. Clearly, uncertainty lies at the heart of statistics and statistical inference. This uncertainty is measured by a probability which therefore forms the crux of statistics and must be properly understood in order to interpret a statistical analysis.

Algorithms↗

Accuracy of MSI testing in predicting germline mutations of MSH2 and MLH1: a case study in Bayesian meta-analysis of diagnostic tests without a gold standard.

Microsatellite instability (MSI) testing is a common screening procedure used to identify families that may harbor mutations of a mismatch repair (MMR) gene and therefore may be at high risk for hereditary colorectal cancer. A reliable estimate of sensitivity and specificity of MSI for detecting germline mutations of MMR genes is critical in genetic counseling and colorectal cancer prevention. Several studies published results of both MSI and mutation analysis on the same subjects. In this article we perform a meta-analysis of these studies and obtain estimates that can be directly used in counseling and screening. In particular, we estimate the sensitivity of MSI for detecting mutations of MSH2 and MLH1 to be 0.81 (0.73-0.89). Statistically, challenges arise from the following: (a) traditional mutation analysis methods used in these studies cannot be considered a gold standard for the identification of mutations; (b) studies are heterogeneous in both the design and the populations considered; and (c) studies may include different patterns of missing data resulting from partial testing of the populations sampled. We address these challenges in the context of a Bayesian meta-analytic implementation of the Hui-Walter design, tailored to account for various forms of incomplete data. Posterior inference is handled via a Gibbs sampler.

Adaptor Proteins, Signal Transducing↗

Advantages of using the net-benefit approach for analysing uncertainty in economic evaluation studies.

No consensus has yet been reached on how to analyse uncertainty in economic evaluation studies where individual patient data are available for costs and health effects. This paper summarises the available results regarding the analysis of uncertainty on the cost-effectiveness plane and argues for using the net-benefit approach when analysing uncertainty in cost-effectiveness studies. The net-benefit approach avoids the interpretation and statistical problems related to the incremental cost effectiveness ratio and implies several advantages. First, traditional statistical methods can be used for confidence-interval estimation and hypothesis testing. Second, calculation of the optimal sample size and the power of the study are facilitated allowing the correlation between costs and effects to vary within and between patient groups. Third, the use of a Bayesian approach to cost-effectiveness analysis is facilitated. Fourth, a formal relation between cost-effectiveness acceptability curves and statistical inference is provided. Finally, the net-benefit approach gives the Fieller's limits of the confidence interval for the incremental cost-effectiveness ratio in the cost-effectiveness plane. Based on these advantages the net-benefit approach should strongly be considered when analysing uncertainty in cost-effectiveness analyses.

Bayes Theorem↗

A Bayesian method for analysing spotted microarray data.

In the decade since their invention, spotted microarrays have been undergoing technical advances that have increased the utility, scope and precision of their ability to measure gene expression. At the same time, more researchers are taking advantage of the fundamentally quantitative nature of these tools with refined experimental designs and sophisticated statistical analyses. These new approaches utilise the power of microarrays to estimate differences in gene expression levels, rather than just categorising genes as up- or down-regulated, and allow the comparison of expression data across multiple samples. In this review, some of the technical aspects of spotted microarrays that can affect statistical inference are highlighted, and a discussion is provided of how several methods for estimating gene expression level across multiple samples deal with these challenges. The focus is on a Bayesian analysis method, BAGEL, which is easy to implement and produces easily interpreted results.

Algorithms↗

Graphical-Model-based Morphometric Analysis.

We propose a novel method for voxel-based morphometry (VBM), which we call Graphical-Model-based Morphometric Analysis (GAMMA), to identify morphological abnormalities automatically, and to find complex probabilistic associations among voxels in magnetic-resonance images and clinical variables. GAMMA is a fully automatic, nonparametric morphometric-analysis algorithm, with high sensitivity and specificity. It uses a Bayesian network to represent the associations among voxels and the function variable, and uses a contextual-clustering method based on a Markov random field to find clusters in which all voxels have similar associations with the function variable. We use loopy belief propagation to infer the unobserved label field and belief map. As opposed to voxel-based morphometric methods based on general linear models, GAMMA is capable of identifying nonlinear associations among the function variable and voxels. Compared with our previous approach, a Bayesian morphometry algorithm, GAMMA has greater sensitivity, specificity, and computational efficiency.

Algorithms↗

Molecular phylogeny of the carnivora (mammalia): assessing the impact of increased sampling on resolving enigmatic relationships.

This study analyzed 76 species of Carnivora using a concatenated sequence of 6243 bp from six genes (nuclear TR-i-I, TBG, and IRBP; mitochondrial ND2, CYTB, and 12S rRNA), representing the most comprehensive sampling yet undertaken for reconstructing the phylogeny of this clade. Maximum parsimony and Bayesian methods were remarkably congruent in topologies observed and in nodal support measures. We recovered all of the higher level carnivoran clades that had been robustly supported in previous analyses (by analyses of morphological and molecular data), including the monophyly of Caniformia, Feliformia, Arctoidea, Pinnipedia, Musteloidea, Procyonidae + Mustelidae sensu stricto, and a clade of (Hyaenidae + (Herpestidae + Malagasy carnivorans)). All of the traditional "families," with the exception of Viverridae and Mustelidae, were robustly supported as monophyletic groups. We further have determined the relative positions of the major lineages within the Caniformia, which previous studies could not resolve, including the first robust support for the phylogenetic position of marine carnivorans (Pinnipedia) within the Arctoidea (as the sister-group to musteloids [sensu lato], with ursids as their sister group). Within the pinnipeds, Odobenidae (walrus) was more closely allied with otariids (sea lions/fur seals) than with phocids ("true" seals). In addition, we recovered a monophyletic clade of skunks and stink badgers (Mephitidae) and resolved the topology of musteloid interrelationships as: Ailurus (Mephitidae (Procyonidae, Mustelidae [sensu stricto])). This pattern of interrelationships of living caniforms suggests a novel inference that large body size may have been the primitive condition for Arctoidea, with secondary size reduction evolving later in some musteloids. Within Mustelidae, Bayesian analyses are unambiguous in supporting otter monophyly (Lutrinae), and in both MP and Bayesian analyses Martes is paraphyletic with respect to Gulo and Eira, as has been observed in some previous molecular studies. Within Feliformia, we have confirmed that Nandinia is the outgroup to all other extant feliforms, and that the Malagasy Carnivora are a monophyletic clade closely allied with the mongooses (Herpestidae [sensu stricto]). Although the monophyly of each of the three major feliform clades (Viverridae sensu stricto, Felidae, and the clade of Hyaenidae + (Herpestidae + Malagasy carnivorans)) is robust in all of our analyses, the relative phylogenetic positions of these three lineages is not resolvable at present. Our analyses document the monophyly of the "social mongooses," strengthening evidence for a single origin of eusociality within the Herpestidae. For a single caniform node, the position of pinnipeds relative to Ursidae and Musteloidea, parsimony analyses of data for the entire Carnivora did not replicate the robust support observed for both parsimony and Bayesian analyses of the caniform ingroup alone. More detailed analyses and these results demonstrate that outgroup choice can have a considerable effect on the strength of support for a particular topology. Therefore, the use of exemplar taxa as proxies for entire clades with diverse evolutionary histories should be approached with caution. The Bayesian analysis likelihood functions generally were better able to reconstruct phylogenetic relationships (increased resolution and more robust support for various nodes) than parsimony analyses when incompletely sampled taxa were included. Bayesian analyses were not immune, however, to the effects of missing data; lower resolution and support in those analyses likely arise from non-overlap of gene sequence data among less well-sampled taxa. These issues are a concern for similar studies, in which different gene sequences are concatenated in an effort to increase resolving power.

Animals↗

Bayesian Gaussian process classification with the EM-EP algorithm.

Gaussian process classifiers (GPCs) are Bayesian probabilistic kernel classifiers. In GPCs, the probability of belonging to a certain class at an input location is monotonically related to the value of some latent function at that location. Starting from a Gaussian process prior over this latent function, data are used to infer both the posterior over the latent function and the values of hyperparameters to determine various aspects of the function. Recently, the expectation propagation (EP) approach has been proposed to infer the posterior over the latent function. Based on this work, we present an approximate EM algorithm, the EM-EP algorithm, to learn both the latent function and the hyperparameters. This algorithm is found to converge in practice and provides an efficient Bayesian framework for learning hyperparameters of the kernel. A multiclass extension of the EM-EP algorithm for GPCs is also derived. In the experimental results, the EM-EP algorithms are as good or better than other methods for GPCs or Support Vector Machines (SVMs) with cross-validation.

Algorithms↗

An ignorant belief network to forecast glucose concentration from clinical databases.

Ignorant Belief Networks (IBNs) are a class of Bayesian Belief Networks (BBNs) able to reason on the basis of incomplete probabilistic information and to incrementally refine the precision of the inferred probabilities as more information becomes available. In this paper, we will describe how can be used to develop a system able to forecast blood glucose concentration in patients affected by insulin dependent diabetes mellitus (IDDM). The major difference between our approach and the traditional ones is that probability distributions over the IBN are not provided by some human expert or by the current literature but they are directly extracted from a clinical database of IDDM patients. This choice capitalizes on the large amount of information generated by the daily control of blood glucose and allows the system to improve the accuracy of predictions as more information becomes available. We will show how, even with a very small subset of the information needed to specify a BBN, the IBN is able to carry out predictions about the future blood glucose concentration in a patient by explicitly taking into consideration the level of ignorance embedded in the network.

Artificial Intelligence↗

Phylogenetic relationships and genotyping of the genus Streptococcus by sequence determination of the RNase P RNA gene, rnpB.

The rnpB gene is universally present in bacterial species and encodes the RNA subunit of endoribonuclease P. In this study, rnpB was sequenced in 50 type strains and 29 additional strains of the genus Streptococcus. Putative secondary-structure models and possible interactions in RNase P RNA molecules are discussed. Phylogenetic relationships were studied and Bayesian, maximum-parsimony and minimum-evolution analyses supported six main clades that comprised 22 of the 50 species analysed. Phylogenetic inference was also studied for the 16S rRNA gene; it indicated a similar tree topology, but with weaker support values than for rnpB. Combined analysis of rnpB and 16S resulted in a phylogeny with significantly better support. Variability in the rnpB and 16S genes among all type strains, calculated as Shannon-Wiener information index values, was 0.45 for rnpB and 0.15 for 16S. Intraspecies proximity was assessed by principal coordinate analysis of rnpB for 32 strains of six closely related species (two clades) and showed species-specific clusters, but heterogeneity occurred in two species. It can be concluded that the rnpB gene is suitable for phylogenetic analysis of closely related taxa and has potential as a tool for species discrimination.

Base Sequence↗

An informative Bayesian structural equation model to assess source-specific health effects of air pollution.

A primary objective of current air pollution research is the assessment of health effects related to specific sources of air particles or particulate matter (PM). Quantifying source-specific risk is a challenge because most PM health studies do not directly observe the contributions of the pollution sources themselves. Instead, given knowledge of the chemical characteristics of known sources, investigators infer pollution source contributions via a source apportionment or multivariate receptor analysis applied to a large number of observed elemental concentrations. Although source apportionment methods are well established for exposure assessment, little work has been done to evaluate the appropriateness of characterizing unobservable sources thus in health effects analyses. In this article, we propose a structural equation framework to assess source-specific health effects using speciated elemental data. This approach corresponds to fitting a receptor model and the health outcome model jointly, such that inferences on the health effects account for the fact that uncertainty is associated with the source contributions. Since the structural equation model (SEM) typically involves a large number of parameters, for small-sample settings, we propose a fully Bayesian estimation approach that leverages historical exposure data from previous related exposure studies. We compare via simulation the performance of our approach in estimating source-specific health effects to that of 2 existing approaches, a tracer approach and a 2-stage approach. Simulation results suggest that the proposed informative Bayesian SEM is effective in eliminating the bias incurred by the 2 existing approaches, even when the number of exposures is limited. We employ the proposed methods in the analysis of a concentrator study investigating the association between ST-segment, a cardiovascular outcome, and major sources of Boston PM and discuss the implications of our findings with respect to the design of future PM concentrator studies.

Air Pollution↗

A phylogenetic framework for the terns (Sternini) inferred from mtDNA sequences: implications for taxonomy and plumage evolution.

We sequenced 2800 bp of mitochondrial DNA from each of 33 species and 2 subspecies (35 taxa) of terns (Sternini), and employed Bayesian methods to derive a phylogeny with good branch support based on posterior probabilities. The resulting tree confirmed many of the generally accepted taxonomic groups, and led us to suggest a revision of the terns that recognizes 12 genera, 11 of which correspond to a distinct clade on the tree or a highly divergent species (1 genus was not represented in the phylogeny). As an example of how the molecular phylogeny reflects similarities in morphology and behavior among the terns, we used the phylogeny to examine the evolution of the breeding (alternate) head plumage patterns among the terns to test the hypothesis that this character is phylogenetically informative. The three basic types of head plumage (white crown, black cap, and black cap with a white blaze on the forehead) were highly conserved within clades, with notable exceptions in two white-crowned species that evolved independently among the black-capped terns. Based on the appearance of the close relatives of these exceptional species, their white crowns appear to be due to the retention of either winter (basic) plumage characteristics or perhaps juvenile characteristics when the birds molt into their breeding plumage. Examination of the evolutionary history of head plumage indicated that the white-crowned species such as the noddies (Anous) and the white tern (Gygis alba) are probably most representative of ancestral terns.

Animals↗

Bayesian analysis of liability of clinical mastitis in Norwegian cattle with a threshold model: effects of data sampling method and model specification.

First-lactation records of Norwegian Cattle were used to infer heritability of liability to clinical mastitis with a threshold sire model. Mastitis was defined as a binary response (presence or absence) in a defined period of first lactation (opportunity period). Length of opportunity period (from 30 d before calving up to 120 or 300 d of lactation) had less effect on heritability estimates than data sampling methods (include or exclude records of cows culled before the end of the opportunity period) whereas sire ranking was more affected by the former. Including all cows, whether culled before the end of the opportunity period or not, gave a sharper and more symmetric posterior distribution of heritability of liability to clinical mastitis. When we analyzed data for all cows, model specification had a small effect on heritability estimates, while sire ranking was affected markedly. Posterior means of heritability range from 0.058 to 0.074. A model regressing on the length of the opportunity period for culled cows without mastitis, was shown favorable for the two opportunity periods using Bayes factors and the deviance information criterion for model comparison. This model, in which liability of mastitis depends on time to culling, may allow utilizing information from all first lactations in genetic evaluation, irrespectively of duration and culling outcome.

Animals↗

Variable selection and Bayesian model averaging in case-control studies.

Covariate and confounder selection in case-control studies is often carried out using a statistical variable selection method, such as a two-step method or a stepwise method in logistic regression. Inference is then carried out conditionally on the selected model, but this ignores the model uncertainty implicit in the variable selection process, and so may underestimate uncertainty about relative risks. We report on a simulation study designed to be similar to actual case-control studies. This shows that p-values computed after variable selection can greatly overstate the strength of conclusions. For example, for our simulated case-control studies with 1000 subjects, of variables declared to be 'significant' with p-values between 0.01 and 0.05, only 49 per cent actually were risk factors when stepwise variable selection was used. We propose Bayesian model averaging as a formal way of taking account of model uncertainty in case-control studies. This yields an easily interpreted summary, the posterior probability that a variable is a risk factor, and our simulation study indicates this to be reasonably well calibrated in the situations simulated. The methods are applied and compared in the context of a case-control study of cervical cancer.

Analysis of Variance↗

Integration of multiple omics reveals key targets and cellular mechanisms for intervention in sarcopenia.

BACKGROUND: Sarcopenia, an age-related syndrome characterized by progressive loss of muscle mass, strength, and function, presents a significant global health burden with limited therapeutic interventions. This study integrates genomic causality, multi-tissue omics, and cellular mediation analyses to identify and prioritize mechanistically grounded therapeutic targets. METHODS: A multi-tiered analytical framework was applied, beginning with two-sample Mendelian randomization (MR) to infer causal relationships between 4907 plasma proteins (cis-pQTLs from 35,559 individuals) and sarcopenia traits in Pan-UK Biobank participants. Bayesian colocalization and transcriptomic validation in human sarcopenia muscle biopsies were employed to prioritize targets. Cellular mediation analysis quantified contributions of immune and stromal cell subtypes to protein-trait pathways using transcriptomic deconvolution. RESULTS: MR identified 1237 plasma proteins causally associated with sarcopenia traits, with six targets (HGFAC, GATM, HMOX2, F2, LMAN2L, HPGDS) validated through colocalization, transcriptomic expression, and sarcopenia-related dysregulation. Cellular mediation revealed immune mechanisms underlying HGFAC's effects, with CD4+ regulatory T cells mediating 3.49 % of its impact on sarcopenia traits. Prothrombin exhibited muscle-protective effects independent of coagulation. CONCLUSION: This study establishes a causal map linking plasma proteins to sarcopenia through immune-stromal interactions. The integration of MR, multi-omics validation, and cellular mediation prioritizes six proteins as actionable targets, supporting repurposing of thrombin inhibitors and development of immunometabolic therapies. The framework bridges genomic causality with cellular pathophysiology, advancing precision strategies for age-related muscle decline.

Humans↗

Mapping mutations on phylogenies.

Mapping of mutations on a phylogeny has been a commonly used analytical tool in phylogenetics and molecular evolution. However, the common approaches for mapping mutations based on parsimony have lacked a solid statistical foundation. Here, I present a Bayesian method for mapping mutations on a phylogeny. I illustrate some of the common problems associated with using parsimony and suggest instead that inferences in molecular evolution can be made on the basis of the posterior distribution of the mappings of mutations. A method for simulating a mapping from the posterior distribution of mappings is also presented, and the utility of the method is illustrated on two previously published data sets. Applications include a method for testing for variation in the substitution rate along the sequence and a method for testing whether the d(N)/d(S) ratio varies among lineages in the phylogeny.

Algorithms↗

A Bayesian missing value estimation method for gene expression profile data.

MOTIVATION: Gene expression profile analyses have been used in numerous studies covering a broad range of areas in biology. When unreliable measurements are excluded, missing values are introduced in gene expression profiles. Although existing multivariate analysis methods have difficulty with the treatment of missing values, this problem has received little attention. There are many options for dealing with missing values, each of which reaches drastically different results. Ignoring missing values is the simplest method and is frequently applied. This approach, however, has its flaws. In this article, we propose an estimation method for missing values, which is based on Bayesian principal component analysis (BPCA). Although the methodology that a probabilistic model and latent variables are estimated simultaneously within the framework of Bayes inference is not new in principle, actual BPCA implementation that makes it possible to estimate arbitrary missing variables is new in terms of statistical methodology. RESULTS: When applied to DNA microarray data from various experimental conditions, the BPCA method exhibited markedly better estimation ability than other recently proposed methods, such as singular value decomposition and K-nearest neighbors. While the estimation performance of existing methods depends on model parameters whose determination is difficult, our BPCA method is free from this difficulty. Accordingly, the BPCA method provides accurate and convenient estimation for missing values. AVAILABILITY: The software is available at http://hawaii.aist-nara.ac.jp/~shige-o/tools/.

Algorithms↗