Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

A fully Bayesian model to cluster gene-expression profiles.

MOTIVATION: With cDNA or oligonucleotide chips, gene-expression levels of essentially all genes in a genome can be simultaneously monitored over a time-course or under different experimental conditions. After proper normalization of the data, genes are often classified into co-expressed classes (clusters) to identify subgroups of genes that share common regulatory elements, a common function or a common cellular origin. With most methods, e.g. k-means, the number of clusters needs to be specified in advance; results depend strongly on this choice. Even with likelihood-based methods, estimation of this number is difficult. Furthermore, missing values often cause problems and lead to the loss of data. RESULTS: We propose a fully probabilistic Bayesian model to cluster gene-expression profiles. The number of classes does not need to be specified in advance; instead it is adjusted dynamically using a Reversible Jump Markov Chain Monte Carlo sampler. Imputation of missing values is integrated into the model. With simulations, we determined the speed of convergence of the sampler as well as the accuracy of the inferred variables. Results were compared with the widely used k-means algorithm. With our method, biologically related co-expressed genes could be identified in a yeast transcriptome dataset, even when some values were missing. AVAILABILITY: The code is available at http://genome.tugraz.at/BayesianClustering/

Algorithms↗

Efficiency of model-based Bayesian methods for detecting hybrid individuals under different hybridization scenarios and with different numbers of loci.

Accurate detection of offspring resulting from hybridization between individuals of distinct populations has a range of applications in conservation and population genetics. We assessed the hybrid identification efficiency of two methods (implemented in the STRUCTURE and NEWHYBRIDS programs) which are tailored to identifying hybrid individuals but use different approaches. Simulated first- and second-generation hybrids were used to assess the performance of these two methods in detecting recent hybridization under scenarios with different levels of genetic divergence and varying numbers of loci. Despite the different approaches of the methods, the hybrid detection efficiency was generally similar and neither of the two methods outperformed the other in all scenarios assessed. Interestingly, hybrid detection efficiency was only minimally affected by whether reference population allele frequency information was included or not. In terms of genotyping effort, efficient detection of F1 hybrid individuals requires the use of 12 or 24 loci with pairwise F(ST) between hybridizing parental populations of 0.21 or 0.12, respectively. While achievable, these locus numbers are nevertheless higher than the number of loci currently commonly applied in population genetic studies. The method of STRUCTURE seemed to be less sensitive to the proportion of hybrids included in the sample, while NEWHYBRIDS seemed to perform slightly better when individuals from both backcross and F1 hybrid classes were present in the sample. However, separating backcrosses from purebred parental individuals requires a considerable genotyping effort (at least 48 loci), even when divergence between parental populations is high.

Animals↗

Bayesian methods for missing covariates in cure rate models.

We propose methods for Bayesian inference for missing covariate data with a novel class of semiparametric survival models with a cure fraction. We allow the missing covariates to be either categorical or continuous and specify a parametric distribution for the covariates that is written as a sequence of one dimensional conditional distributions. We assume that the missing covariates are missing at random (MAR) throughout. We propose an informative class of joint prior distributions for the regression coefficients and the parameters arising from the covariate distributions. The proposed class of priors are shown to be useful in recovering information on the missing covariates especially in situations where the missing data fraction is large. Properties of the proposed prior and resulting posterior distributions are examined. Also, model checking techniques are proposed for sensitivity analyses and for checking the goodness of fit of a particular model. Specifically, we extend the Conditional Predictive Ordinate (CPO) statistic to assess goodness of fit in the presence of missing covariate data. Computational techniques using the Gibbs sampler are implemented. A real data set involving a melanoma cancer clinical trial is examined to demonstrate the methodology.

Bayes Theorem↗

Model-independent mean-field theory as a local method for approximate propagation of information.

We present a systematic approach to mean-field theory (MFT) in a general probabilistic setting without assuming a particular model. The mean-field equations derived here may serve as a local, and thus very simple, method for approximate inference in probabilistic models such as Boltzmann machines or Bayesian networks. Our approach is 'model-independent' in the sense that we do not assume a particular type of dependences; in a Bayesian network, for example, we allow arbitrary tables to specify conditional dependences. In general, there are multiple solutions to the mean-field equations. We show that improved estimates can be obtained by forming a weighted mixture of the multiple mean-field solutions. Simple approximate expressions for the mixture weights are given. The general formalism derived so far is evaluated for the special case of Bayesian networks. The benefits of taking into account multiple solutions are demonstrated by using MFT for inference in a small and in a very large Bayesian network. The results are compared with the exact results.

Child↗

Approaches for optimal sequential decision analysis in clinical trials.

Unlike traditional approaches, Bayesian methods enable formal combination of expert opinion and objective information into interim and final analyses of clinical trial data. However, most previous Bayesian approaches have based the stopping decision on the posterior probability content of one or more regions of the parameter space, thus implicitly determining a loss and decision structure. In this paper, we offer a fully Bayesian approach to this problem, specifying not only the likelihood and prior distributions but appropriate loss functions as well. At each data monitoring point, we enumerate the available decisions and investigate the use of backward induction, implemented via Monte Carlo methods, to choose the optimal course of action. We then present a forward sampling algorithm that substantially eases the analytic and computational burdens associated with backward induction, offering the possibility of fully Bayesian optimal sequential monitoring for previously untenable numbers of interim looks. We show that forward sampling can always identify the optimal sequential strategy in the case of a one-parameter exponential family with a conjugate prior and monotone loss functions as well as the best member of a certain class of strategies when backward induction is infeasible. Finally, we illustrate and compare the forward and backward approaches using data from a recent AIDS clinical trial.

AIDS-Related Opportunistic Infections↗

Inferring the root of a phylogenetic tree.

Phylogenetic trees can be rooted by a number of criteria. Here, we introduce a Bayesian method for inferring the root of a phylogenetic tree by using one of several criteria: the outgroup, molecular clock, and nonreversible model of DNA substitution. We perform simulation analyses to examine the relative ability of these three criteria to correctly identify the root of the tree. The outgroup and molecular clock criteria were best able to identify the root of the tree, whereas the nonreversible model was able to identify the root only when the substitution process was highly nonreversible. We also examined the performance of the criteria for a tree of four species for which the topology and root position are well supported. Results of the analyses of these data are consistent with the simulation results.

Bayes Theorem↗

Tomographic reconstruction using 3D deformable models.

We address the issue of reconstructing the shape of an object with uniform interior activity from a set of projections. We estimate directly from projection data the position of a triangulated surface describing the boundary of the object while incorporating prior knowledge about the unknown shape. This inverse problem is addressed in a Bayesian framework using the maximum a posteriori (MAP) estimate for the reconstruction. The derivatives needed for the gradient-based optimization of the model parameters are obtained using the adjoint differentiation technique. We present results from a numerical simulation of a dynamic cardiac imaging study. A first-pass exam is simulated with a numerical phantom of the right ventricle using the measured system response of the University of Arizona FASTSPECT imager, which consists of 24 detectors. We demonstrate the usefulness of our approach by reconstructing the shape of the ventricle from 10,000 counts. The comparison with an ML-EM result shows the usefulness of the deformable model approach.

Bayes Theorem↗

Predicting event times in clinical trials when treatment arm is masked.

Because power is primarily determined by the number of events in event-based clinical trials, the timing for interim or final analysis of data is often determined based on the accrual of events during the course of the study. Thus, it is of interest to predict early and accurately the time of a landmark interim or terminating event. Existing Bayesian methods may be used to predict the date of the landmark event, based on current enrollment, event, and loss to follow-up, if treatment arms are known. This work extends these methods to the case where the treatment arms are masked by using a parametric mixture model with a known mixture proportion. Posterior simulation using the mixture model is compared with methods assuming a single population. Comparison of the mixture model with the single-population approach shows that with few events, these approaches produce substantially different results and that these results converge as the prediction time is closer to the landmark event. Simulations show that the mixture model with diffuse priors can have better coverage probabilities for the prediction interval than the nonmixture models if a treatment effect is present.

Bayes Theorem↗

Using fragment chemistry data mining and probabilistic neural networks in screening chemicals for acute toxicity to the fathead minnow.

The paper is illustrating how the general data mining methodology may be adapted to provide solutions to the problem of high throughput virtual screening of organic chemicals for possible acute toxicity to the fathead minnow fish. The present approach involves mining fragment information from chemical structures and is using probabilistic neural networks to model the relationship between structure and toxicity. Probabilistic neural networks implement a special class of multivariate non-linear Bayesian statistical models. The mathematical principles supporting their use for value prediction purposes are clarified and their peculiarities discussed. As part of the research phase of the data mining process, a dataset consisting of 800 structures and associated fathead minnow (Pimephales promelas) 96-h LC50 acute toxicity endpoint information is used for both the purpose of identifying an advantageous combination of fragment descriptors and for training the neural networks. As a result, two powerful models are generated. Model 1 implements the basic PNN with Gaussian kernel (statistical corrections included) while Model 2 implements the PNN with Gaussian kernel and separated variables. External validation is performed using a separate dataset consisting of 86 structures and associated toxicity information. Both learning and generalization capabilities of the two models are investigated and their limitations discussed.

Animals↗

A Bayesian space varying parameter model applied to estimating fertility schedules.

We propose a spatial generalized linear model (GLM) to analyse the vital rates for small areas. In each small area, we have a response vector and covariates to explain its variability. The statistical methodology is based on a spatial Bayesian approach and it allows the covariates' parameters of the generalized linear model to vary smoothly on space. Hence, the effect of a covariate on the response varies depending on the random variables measurement location. Our model is an extension of disease mapping models allowing the space-covariate interaction to be modelled in a natural way and giving space a position of intrinsic interest. We introduce the model in the context of fertility curve estimation. In each small area, we have a curve describing the variation of fertility rates by age modelled by Coale's fertility model, which implies a GLM in each area. A simulation shows the advantages of our approach. In addition, the paper applies the procedure to census data used to study the diffusion of low fertility behaviour in Brazil.

Adolescent↗

A Bayesian compound stochastic process for modeling nonstationary and nonhomogeneous sequence evolution.

Variations of nucleotidic composition affect phylogenetic inference conducted under stationary models of evolution. In particular, they may cause unrelated taxa sharing similar base composition to be grouped together in the resulting phylogeny. To address this problem, we developed a nonstationary and nonhomogeneous model accounting for compositional biases. Unlike previous nonstationary models, which are branchwise, that is, assume that base composition only changes at the nodes of the tree, in our model, the process of compositional drift is totally uncoupled from the speciation events. In addition, the total number of events of compositional drift distributed across the tree is directly inferred from the data. We implemented the method in a Bayesian framework, relying on Markov Chain Monte Carlo algorithms, and applied it to several nucleotidic data sets. In most cases, the stationarity assumption was rejected in favor of our nonstationary model. In addition, we show that our method is able to resolve a well-known artifact. By Bayes factor evaluation, we compared our model with 2 previously developed nonstationary models. We show that the coupling between speciations and compositional shifts inherent to branchwise models may lead to an overparameterization, resulting in a lesser fit. In some cases, this leads to incorrect conclusions, concerning the nature of the compositional biases. In contrast, our compound model more flexibly adapts its effective number of parameters to the data sets under investigation. Altogether, our results show that accounting for nonstationary sequence evolution may require more elaborate and more flexible models than those currently used.

Animals↗

Bayesian optimal designs for a quantal dose-response study with potentially missing observations.

In a dose-response study, there are frequently multiple goals and not all planned observations are realized at the end of the study. Subjects drop out and the initial design can be quite different from the final design. Consequently, the final design can be inefficient. Single- and multiple-objective Bayesian optimal designs that account for potentially missing observations in quantal response models were recently proposed in Baek (2005). In this work, we investigate the efficiencies of the conventional optimal designs that do not incorporate potential missing information relative to our proposed designs. Furthermore, we examine the impact of restricted dose range on the resulting optimal designs. As an application, we used missing data information from a study by Yocum et al. (2003) to design a study for estimating dose levels of tacrolimus that will result in a certain percentage of rheumatoid arthritis patients having an ACR20 response at 6 months.

Algorithms↗

Gradient descent learning in and out of equilibrium.

Relations between the off thermal equilibrium dynamical process of on-line learning and the thermally equilibrated off-line learning are studied for potential gradient descent learning. The approach of Opper to study on-line Bayesian algorithms is used for potential based or maximum likelihood learning. We look at the on-line learning algorithm that best approximates the off-line algorithm in the sense of least Kullback-Leibler information loss. The closest on-line algorithm works by updating the weights along the gradient of an effective potential, which is different from the parent off-line potential. A few examples are analyzed and the origin of the potential annealing is discussed.

Algorithms↗

Generalization of map estimation in SAAM II: validation against ADAPT II in a glucose model case study.

Bayesian approaches to model identification [e.g., maximum a posteriori (MAP) estimation] are receiving increasing attention in metabolism since important quantitative knowledge has become available in the last decades, e.g., from tracer experiments. By suitably exploiting this knowledge, more complex physiological models than those solely based on experimental data (Fisherian approach) become resolvable. While ADAPT II is the reference software for MAP estimation in pharmacokinetic/pharmacodynamic/metabolic system analysis, another popular, user-friendly and state-of-the-art software is SAAM II. However, SAAM II does not handle a priori information on correlation among parameters, thus allowing a limited version of MAP estimation to be performed. The aim here is twofold. First, we show that this limitation of SAAM II can be easily overcome by resorting to a probability theory result. Second, we test SAAM II vs ADAPT II implementation of MAP estimation in a real case study: the Bayesian identification of a recently proposed two-compartment minimal model of glucose kinetics during an intravenous glucose tolerance test. SAAM II MAP estimates of glucose effectiveness (SG) and insulin sensitivity (S(I)) obtained in a group of 22 healthy humans are in excellent agreement with those of ADAPT II: S(G) = 2.84 +/- 0.27 vs. 2.84 +/- 0.27 (mlmin(-1) kg(-1), mean +/- SD) and S(I) = 11.46 +/- 1.69 vs. 11.47 +/- 1.69 [10(-2) ml kg(-1) min(-1)/ (microU ml(-1))]. The SAAM II vs. ADAPT II estimates are virtually identical (P > 0.44 and 0.68 for S(G) and S(I), respectively) and also closely correlated (p = 0.9998 and 0.9999).

Algorithms↗

Model selection for mixtures of mutagenetic trees.

The evolution of drug resistance in HIV is characterized by the accumulation of resistance-associated mutations in the HIV genome. Mutagenetic trees, a family of restricted Bayesian tree models, have been applied to infer the order and rate of occurrence of these mutations. Understanding and predicting this evolutionary process is an important prerequisite for the rational design of antiretroviral therapies. In practice, mixtures models of K mutagenetic trees provide more flexibility and are often more appropriate for modelling observed mutational patterns. Here, we investigate the model selection problem for K-mutagenetic trees mixture models. We evaluate several classical model selection criteria including cross-validation, the Bayesian Information Criterion (BIC), and the Akaike Information Criterion. We also use the empirical Bayes method by constructing a prior probability distribution for the parameters of a mutagenetic trees mixture model and deriving the posterior probability of the model. In addition to the model dimension, we consider the redundancy of a mixture model, which is measured by comparing the topologies of trees within a mixture model. Based on the redundancy, we propose a new model selection criterion, which is a modification of the BIC. Experimental results on simulated and on real HIV data show that the classical criteria tend to select models with far too many tree components. Only cross-validation and the modified BIC recover the correct number of trees and the tree topologies most of the time. At the same optimal performance, the runtime of the new BIC modification is about one order of magnitude lower. Thus, this model selection criterion can also be used for large data sets for which cross-validation becomes computationally infeasible.

Bayes Theorem↗

Unified SPM-ICA for fMRI analysis.

A widely used tool for functional magnetic resonance imaging (fMRI) data analysis, statistical parametric mapping (SPM), is based on the general linear model (GLM). SPM therefore requires a priori knowledge or specific assumptions about the time courses contributing to signal changes. In contradistinction, independent component analysis (ICA) is a data-driven method based on the assumption that the causes of responses are statistically independent. Here we describe a unified method, which combines ICA, temporal ICA (tICA), and SPM for analyzing fMRI data. tICA was applied to fMRI datasets to disclose independent components, whose number was determined by the Bayesian information criterion (BIC). The resulting components were used to construct the design matrix of a GLM. Parameters were estimated and regionally-specific statistical inferences were made about activations in the usual way. The sensitivity and specificity were evaluated using Monte Carlo simulations. The receiver operating characteristic (ROC) curves indicated that the unified SPM-ICA method had a better performance. Moreover, SPM-ICA was applied to fMRI datasets from twelve normal subjects performing left and right hand movements. The areas identified corresponded to motor (premotor, sensorimotor areas and SMA) areas and were consistently task related. Part of the frontal lobe, parietal cortex, and cingulate gyrus also showed transiently task-related responses. The unified method requires less supervision than the conventional SPM and enables classical inference about the expression of independent components. Our results also suggest that the method has a higher sensitivity than SPM analyses.

Adult↗

2D Autocorrelation modeling of the negative inotropic activity of calcium entry blockers using Bayesian-regularized genetic neural networks.

Negative inotropic potency of 60 benzothiazepine-like calcium entry blockers (CEBs), Diltiazem analogs, was successfully modeled using Bayesian-regularized genetic neural networks (BRGNNs) and 2D autocorrelation vectors. This approach yielded reliable and robust models whilst by means of a linear genetic algorithm (GA) search routine no multilinear regression model was found describing more than 50% of the training set. On the contrary, the optimum neural network predictor with five inputs described about 84% and 65% variances of 50 randomly selected training and test sets. Autocorrelation vectors in the nonlinear model contained information regarding 2D spatial distributions on the CEB structure of van der Waals volumes, electronegativities, and polarizabilities. However, a sensitivity analysis of the network inputs pointed out to the electronegativity and polarizability 2D topological distributions at substructural fragments of sizes 3 and 4 as the most relevant features governing the nonlinear modeling of the negative inotropic potency.

Bayes Theorem↗

Simultaneous determination of thiocyanate and salicylate by a combined UV-spectrophotometric detection principal component artificial neural network.

A modified principle component artificial neural network (PC-ANN) model is developed for simultaneous determination of thiocyanate and salycilate concentration after passing through the bulk of a liquid membrane by tri-phenyl benzyl phosphonium chloride. All calibration, and test samples data were obtained using UV-Vis spectrophotometer. In this way, a modified PC-ANN consisting of three layers of nodes was trained by combination of Bayesian-Levenberg-Marquardt as training rule. Sigmoid and liner transfer functions were used in the hidden and output layers respectively to facilitate nonlinear calibration. The model could accurately estimate the concentration of components with acceptable precision and accuracy, for mixtures. The PC-ANN model exhibits a good ability for the simultaneous determination of the thiocyanate and salycilate in concentration range 0.5 x 10(-4) mol.l(-1) up to 5.0 x 10(-4) mol.l(-1) with Root Mean square error (2.22% and 2.20%, for thiocyanate and salycilate, respectively) and high correlation coefficients (R2= 0.998 or greater). Results obtained with modified trained PC-ANN were compared with stepwise linear regression (SMLR) model. Validation of the two models shows a better ability in estimation of the modified PC-ANN as compared with the SMLR model (MSRE given are 3.12%, 6.31%.).

Neural Networks, Computer↗