Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Comparison of methods for analysing cluster randomized trials: an example involving a factorial design.

BACKGROUND: Studies involving clustering effects are common, but there is little consistency in their analysis. Various analytical methods were compared for a factorial cluster randomized trial (CRT) of two primary care-based interventions designed to increase breast screening attendance. METHODS: Three cluster-level and five individual-level options were compared in respect of log odds ratios of attendance and their standard errors (SE), for the two intervention effects and their interaction. Cluster-level analyses comprised: (C1) unweighted regression of practice log odds; (C2) regression of log odds weighted by their inverse variance; (C3) random-effects meta-regression of log odds with practice as a random effect. Individual-level analyses comprised: (I1) standard logistic regression ignoring clustering; (I2) robust SE; (I3) generalized estimating equations; (I4) random-effects logistic regression; (I5) Bayesian random-effects logistic regression. Adjustments for stratification and baseline variables were investigated. RESULTS: As expected, method I1 was highly anti-conservative. The other, valid, methods exhibited considerable differences in parameter estimates and standard errors, even between the various random-effects methods based on the same statistical model. Method I4 was particularly sensitive to between-cluster variation and was computationally stable only after controlling for baseline uptake. CONCLUSIONS: Commonly used methods for the analysis of CRT can give divergent results. Simulation studies are needed to compare results from different methods in situations typical of cluster trials but when the true model parameters are known.

Breast Neoplasms↗

Bayesian identification of admixture events using multilocus molecular markers.

Bayesian statistical methods for the estimation of hidden genetic structure of populations have gained considerable popularity in the recent years. Utilizing molecular marker data, Bayesian mixture models attempt to identify a hidden population structure by clustering individuals into genetically divergent groups, whereas admixture models target at separating the ancestral sources of the alleles observed in different individuals. We discuss the difficulties involved in the simultaneous estimation of the number of ancestral populations and the levels of admixture in studied individuals' genomes. To resolve this issue, we introduce a computationally efficient method for the identification of admixture events in the population history. Our approach is illustrated by analyses of several challenging real and simulated data sets. The software (baps), implementing the methods introduced here, is freely available at http://www.rni.helsinki.fi/~jic/bapspage.html.

Alleles↗

Assessing the accuracy of ancestral protein reconstruction methods.

The phylogenetic inference of ancestral protein sequences is a powerful technique for the study of molecular evolution, but any conclusions drawn from such studies are only as good as the accuracy of the reconstruction method. Every inference method leads to errors in the ancestral protein sequence, resulting in potentially misleading estimates of the ancestral protein's properties. To assess the accuracy of ancestral protein reconstruction methods, we performed computational population evolution simulations featuring near-neutral evolution under purifying selection, speciation, and divergence using an off-lattice protein model where fitness depends on the ability to be stable in a specified target structure. We were thus able to compare the thermodynamic properties of the true ancestral sequences with the properties of "ancestral sequences" inferred by maximum parsimony, maximum likelihood, and Bayesian methods. Surprisingly, we found that methods such as maximum parsimony and maximum likelihood that reconstruct a "best guess" amino acid at each position overestimate thermostability, while a Bayesian method that sometimes chooses less-probable residues from the posterior probability distribution does not. Maximum likelihood and maximum parsimony apparently tend to eliminate variants at a position that are slightly detrimental to structural stability simply because such detrimental variants are less frequent. Other properties of ancestral proteins might be similarly overestimated. This suggests that ancestral reconstruction studies require greater care to come to credible conclusions regarding functional evolution. Inferred functional patterns that mimic reconstruction bias should be reevaluated.

Algorithms↗

Biochemical networks with uncertain parameters.

The modelling of biochemical networks becomes delicate if kinetic parameters are varying, uncertain or unknown. Facing this situation, we quantify uncertain knowledge or beliefs about parameters by probability distributions. We show how parameter distributions can be used to infer probabilistic statements about dynamic network properties, such as steady-state fluxes and concentrations, signal characteristics or control coefficients. The parameter distributions can also serve as priors in Bayesian statistical analysis. We propose a graphical scheme, the 'dependence graph', to bring out known dependencies between parameters, for instance, due to the equilibrium constants. If a parameter distribution is narrow, the resulting distribution of the variables can be computed by expanding them around a set of mean parameter values. We compute the distributions of concentrations, fluxes and probabilities for qualitative variables such as flux directions. The probabilistic framework allows the study of metabolic correlations, and it provides simple measures of variability and stochastic sensitivity. It also shows clearly how the variability of biological systems is related to the metabolic response coefficients.

Animals↗

Bayesian methods for estimating pathogen prevalence within groups of animals from faecal-pat sampling.

Pathogens such as Escherichia coli O157:H7 and Campylobacter spp. have been implicated in outbreaks of food poisoning in the UK and elsewhere. Domestic animals and wildlife are important reservoirs for both of these agents, and cross-contamination from faeces is believed to be responsible for many human outbreaks. Appropriate parameterisation of quantitative microbial-risk models requires representative data at all levels of the food chain. Our focus in this paper is on the early stages of the food chain-specifically, sampling issues which arise at the farm level. We estimated animal-pathogen prevalence from faecal-pat samples using a Bayesian method which reflected the uncertainties inherent in the animal-level prevalence estimates. (Note that prevalence here refers to the percentage of animals shedding the bacteria of interest). The method offers more flexibility than traditional, classical approaches: it allows the incorporation of prior belief, and permits the computation of a variety of distributional and numerical summaries, analogues of which often are not available through a classical framework. The Bayesian technique is illustrated with a number of examples reflecting the effects of a diversity of assumptions about the underlying processes. The technique appears to be both robust and flexible, and is useful when defecation rates in infected and uninfected groups are unequal, where population size is uncertain, and also where the microbiological-test sensitivity is imperfect. We also investigated the determination of the sample size necessary for determining animal-level prevalence from pat samples to within a pre-specified degree of accuracy.

Agriculture↗

A tracking approach to parcellation of the cerebral cortex.

The cerebral cortex is composed of regions with distinct laminar structure. Functional neuroimaging results are often reported with respect to these regions, usually by means of a brain "atlas". Motivated by the need for more precise atlases, and the lack of model-based approaches in prior work in the field, this paper introduces a novel approach to parcellating the cortex into regions of distinct laminar structure, based on the theory of target tracking. The cortical layers are modelled by hidden Markov models and are tracked to determine the Bayesian evidence of layer hypotheses. This model-based parcellation method, evaluated here on a set of histological images of the cortex, is extensible to 3-D images.

Algorithms↗

Inferential backbone assignment for sparse data.

This paper develops an approach to protein backbone NMR assignment that effectively assigns large proteins while using limited sets of triple-resonance experiments. Our approach handles proteins with large fractions of missing data and many ambiguous pairs of pseudoresidues, and provides a statistical assessment of confidence in global and position-specific assignments. The approach is tested on an extensive set of experimental and synthetic data of up to 723 residues, with match tolerances of up to 0.5 ppm for Calpha and Cbeta resonance types. The tests show that the approach is particularly helpful when data contain experimental noise and require large match tolerances. The keys to the approach are an empirical Bayesian probability model that rigorously accounts for uncertainty in the data at all stages in the analysis, and a hybrid stochastic tree-based search algorithm that effectively explores the large space of possible assignments.

Algorithms↗

A heuristic Bayesian method for segmenting DNA sequence alignments and detecting evidence for recombination and gene conversion.

We propose a heuristic approach to the detection of evidence for recombination and gene conversion in multiple DNA sequence alignments. The proposed method consists of two stages. In the first stage, a sliding window is moved along the DNA sequence alignment, and phylogenetic trees are sampled from the conditional posterior distribution with MCMC. To reduce the noise intrinsic to inference from the limited amount of data available in the typically short sliding window, a clustering algorithm based on the Robinson-Foulds distance is applied to the trees thus sampled, and the posterior distribution over tree clusters is obtained for each window position. While changes in this posterior distribution are indicative of recombination or gene conversion events, it is difficult to decide when such a change is statistically significant. This problem is addressed in the second stage of the proposed algorithm, where the distributions obtained in the first stage are post-processed with a Bayesian hidden Markov model (HMM). The emission states of the HMM are associated with posterior distributions over phylogenetic tree topology clusters. The hidden states of the HMM indicate putative recombinant segments. Inference is done in a Bayesian sense, sampling parameters from the posterior distribution with MCMC. Of particular interest is the determination of the number of hidden states as an indication of the number of putative recombinant regions. To this end, we apply reversible jump MCMC, and sample the number of hidden states from the respective posterior distribution.

Actins↗

Analysis of planar shapes using geodesic paths on shape spaces.

For analyzing shapes of planar, closed curves, we propose differential geometric representations of curves using their direction functions and curvature functions. Shapes are represented as elements of infinite-dimensional spaces and their pairwise differences are quantified using the lengths of geodesics connecting them on these spaces. We use a Fourier basis to represent tangents to the shape spaces and then use a gradient-based shooting method to solve for the tangent that connects any two shapes via a geodesic. Using the Surrey fish database, we demonstrate some applications of this approach: 1) interpolation and extrapolations of shape changes, 2) clustering of objects according to their shapes, 3) statistics on shape spaces, and 4) Bayesian extraction of shapes in low-quality images.

Algorithms↗

Genetic change for clinical mastitis in Norwegian cattle: a threshold model analysis.

Records of clinical mastitis on 1.6 million first-lactation daughters of 2,411 Norwegian Cattle sires that were progeny tested from 1978 through 1998 were analyzed with a threshold model. The main objective was to infer genetic change for the disease in the population. A Bayesian approach via Gibbs sampling was used. The model for the underlying liability had age at first calving, month x year of calving, herd x 3-year-period, and sire of the cow as explanatory variables. Posterior mean (SD) of heritability of liability to clinical mastitis was 0.066 (0.003). Genetic evaluations (posterior means) of sires both in the liability and observable scales were computed. Annual genetic change of liability to clinical mastitis for progeny tested bulls born from 1973 to 1993 was assessed. The linear regression of mean sire effect on year of birth had a posterior mean (SD) of -0.00018 (0.0004), suggesting a nearly constant genetic level for clinical mastitis. However, an analysis of sire posterior means by birth-year of daughters indicated an approximately constant genetic level in the cow population from 1976 to 1990 (-0.02%/yr), and a genetic improvement thereafter (-0.27%/yr). This reflects more emphasis on mastitis in selection of bulls in recent years. Corresponding results obtained with a standard linear model analysis were -0.01% and -0.23% per year, respectively (regression of sire predicted transmitting ability on birth-year of daughters). Genetic change seems to be slightly understated with the linear model, assuming the threshold model holds true.

Animals↗

Bayesian diagnostic probabilities without assuming independence of symptoms.

The paper describes an application of Bayes' Theorem to the problem of estimating from past data the probabilities that patients have certain diseases, given their symptoms. The data consist of hospital records of patients who suffered acute abdominal pain. For each patient the records showed a large number of symptoms and the final diagnosis to one of nine diseases or diagnostic groups. Most current methods of computer diagnosis use the "Simple Bayes" model in which the symptoms are assumed to be independent, but the present paper does not make this assumption. Those symptoms (or lack of symptoms) which are most relevant to the diagnosis of each disease are identified by a sequence of chi-squared tests. The computer diagnoses obtained as a result of the implementation of this approach are compared with those given by the "Simple Bayes" method, by the method of classification trees (CART), and also with the preliminary and final diagnoses made by physicians.

Abdominal Pain↗

Bayesian methods for pharmacokinetic models in dynamic contrast-enhanced magnetic resonance imaging.

This paper proposes a new method for estimating kinetic parameters of dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) based on adaptive Gaussian Markov random fields. Kinetic parameter estimates using neighboring voxels reduce the observed variability in local tumor regions while preserving sharp transitions between heterogeneous tissue boundaries. Asymptotic results for standard errors from likelihood-based nonlinear regression are compared with those derived from the posterior distribution using Bayesian estimation with and without neighborhood information. Application of the method to the analysis of breast tumors based on kinetic parameters has shown that the use of Bayesian analysis combined with adaptive Gaussian Markov random fields provides improved convergence behavior and more consistent morphological and functional statistics.

Algorithms↗

Probability of a segregating pattern in a sample of DNA sequences.

Mutations that result in segregating sites (polymorphic sites) in a sample of DNA sequences can be classified into different types. A pattern of segregating sites is an array of the numbers of various types of mutations. Using an urn model, the probability of a pattern of segregating sites can be expressed as a recurrence equation and its value can be computed sequentially. Among those that can be computed by this method are the probability of obtaining k external mutations (mutations that occur in external branches of the genealogy of a sample), the probability of obtaining k internal mutations (mutations that occur in internal branches), the probability of obtaining k singletons (segregating sites at which one of the two segregating nucleotides is present in only one sequence), and the probability of obtaining k non-singletons. Two applications of the method are discussed. One is a maximum likelihood estimation of straight theta and another is a Bayesian statistical test of the hypothesis of neutral mutations.

Bayes Theorem↗

Bayesian approach for neural networks--review and case studies.

We give a short review on the Bayesian approach for neural network learning and demonstrate the advantages of the approach in three real applications. We discuss the Bayesian approach with emphasis on the role of prior knowledge in Bayesian models and in classical error minimization approaches. The generalization capability of a statistical model, classical or Bayesian, is ultimately based on the prior assumptions. The Bayesian approach permits propagation of uncertainty in quantities which are unknown to other assumptions in the model, which may be more generally valid or easier to guess in the problem. The case problem studied in this paper include a regression, a classification, and an inverse problem. In the most thoroughly analyzed regression problem, the best models were those with less restrictive priors. This emphasizes the major advantage of the Bayesian approach, that we are not forced to guess attributes that are unknown, such as the number of degrees of freedom in the model, non-linearity of the model with respect to each input variable, or the exact form for the distribution of the model residuals.

Algorithms↗

Bayesian tests of extra-Binomial variability.

A simple model for extra-Binomial variability is the Beta-Binomial. A complication in testing the Binomial against the Beta-Binomial alternative is that the Binomial lies on the boundary of the Beta-Binomial, which forces modifications to the usual asymptotic arguments. In this paper, we propose a Bayesian test using a pair of approximate Bayes factors, one for the case in which the maximum likelihood estimator (MLE) of the extra-Binomial variability is zero and one for the case in which it is positive. These approximate Bayes factors are easy to compute. We evaluate the operating characteristics of the Bayes factors and find them to be more powerful than the likelihood ratio test. We then apply the method to three data sets, including one in which the issue is whether a logistic regression intercept should be considered a random effect. In each case, our approximate Bayes factors are close to the exact Bayes factors, which may also be computed with additional effort.

Amputation, Surgical↗

Mixed model analysis of quantitative trait loci.

We develop a mixed model approach of quantitative trait locus (QTL) mapping for a hybrid population derived from the crosses of two or more distinguished outbred populations. Under the mixed model, we treat the mean allelic value of each source population as the fixed effect and the allelic deviations from the mean as random effects so that we can partition the total genetic variance into between- and within-population variances. Statistical inference of the QTL parameters is obtained by using the Bayesian method implemented by Markov chain Monte Carlo (MCMC). This unified QTL mapping algorithm treats the fixed and random model approaches as special cases of the general mixed model methodology. Utility and flexibility of the method are demonstrated by using a set of simulated data.

Algorithms↗

Finding optimal models for small gene networks.

Finding gene networks from microarray data has been one focus of research in recent years. Given search spaces of super-exponential size, researchers have been applying heuristic approaches like greedy algorithms or simulated annealing to infer such networks. However, the accuracy of heuristics is uncertain, which--in combination with the high measurement noise of microarrays--makes it very difficult to draw conclusions from networks estimated by heuristics. We present a method that finds optimal Bayesian networks of considerable size and show first results of the application to yeast data. Having removed the uncertainty due to the heuristic methods, it becomes possible to evaluate the power of different statistical models to find biologically accurate networks.

Algorithms↗

Evaluating the quality of a probabilistic diagnostic system using different inferencing strategies.

In this paper we describe the evaluation of a probabilistic diagnostic system for patients with renal mass. Three inference models: Multi-membership Bayesian (MB), Minimal Diagnosis (MD) and Bayesian Network (BN), and 72 patients are used to illustrate three interrelated measures of system performance: accuracy, reliability and discriminating power. The inferencing strategies we tested demonstrated the kind of trade-offs in the performance measures that can be expected from imperfect systems. Ultimately, the purpose and expected use of a system should dictate the relative importance ascribed to different aspects of system performance.

Adolescent↗