Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Improving Cox survival analysis with a neural-Bayesian approach.

In this article we show that traditional Cox survival analysis can be improved upon when supplemented with sensible priors and analysed within a neural Bayesian framework. We demonstrate that the Bayesian method gives more reliable predictions, in particular for relatively small data sets. The obtained posterior (the probability distribution of network parameters given the data) which in itself is intractable, can be made accessible by several approximations. We review approximations by Hybrid Markov Chain Monte Carlo sampling, a variational method and the Laplace approximation. We argue that although each Bayesian approach circumvents the shortcomings of the original Cox analysis, and therefore yields better predictive results, in practice the use of variational methods or Laplace is preferable. Since Cox survival analysis is infamous for its poor results with (too) many inputs, we use the Bayesian posterior to estimate p-values on the inputs and to formulate an algorithm for backward elimination. We show that after removal of irrelevant inputs Bayesian methods still achieve significantly better results than classical Cox.

Antineoplastic Agents↗

A Bayesian method for analysing spotted microarray data.

In the decade since their invention, spotted microarrays have been undergoing technical advances that have increased the utility, scope and precision of their ability to measure gene expression. At the same time, more researchers are taking advantage of the fundamentally quantitative nature of these tools with refined experimental designs and sophisticated statistical analyses. These new approaches utilise the power of microarrays to estimate differences in gene expression levels, rather than just categorising genes as up- or down-regulated, and allow the comparison of expression data across multiple samples. In this review, some of the technical aspects of spotted microarrays that can affect statistical inference are highlighted, and a discussion is provided of how several methods for estimating gene expression level across multiple samples deal with these challenges. The focus is on a Bayesian analysis method, BAGEL, which is easy to implement and produces easily interpreted results.

Algorithms↗

Predicting protein secondary structure with probabilistic schemata of evolutionarily derived information.

We demonstrate the applicability of our previously developed Bayesian probabilistic approach for predicting residue solvent accessibility to the problem of predicting secondary structure. Using only single-sequence data, this method achieves a three-state accuracy of 67% over a database of 473 non-homologous proteins. This approach is more amenable to inspection and less likely to overlearn specifics of a dataset than "black box" methods such as neural networks. It is also conceptually simpler and less computationally costly. We also introduce a novel method for representing and incorporating multiple-sequence alignment information within the prediction algorithm, achieving 72% accuracy over a dataset of 304 non-homologous proteins. This is accomplished by creating a statistical model of the evolutionarily derived correlations between patterns of amino acid substitution and local protein structure. This model consists of parameter vectors, termed "substitution schemata," which probabilistically encode the structure-based heterogeneity in the distributions of amino acid substitutions found in alignments of homologous proteins. The model is optimized for structure prediction by maximizing the mutual information between the set of schemata and the database of secondary structures. Unlike "expert heuristic" methods, this approach has been demonstrated to work well over large datasets. Unlike the opaque neural network algorithms, this approach is physicochemically intelligible. Moreover, the model optimization procedure, the formalism for predicting one-dimensional structural features and our previously developed method for tertiary structure recognition all share a common Bayesian probabilistic basis. This consistency starkly contrasts with the hybrid and ad hoc nature of methods that have dominated this field in recent years.

Algorithms↗

Modifying the Schwarz Bayesian information criterion to locate multiple interacting quantitative trait loci.

The problem of locating multiple interacting quantitative trait loci (QTL) can be addressed as a multiple regression problem, with marker genotypes being the regressor variables. An important and difficult part in fitting such a regression model is the estimation of the QTL number and respective interactions. Among the many model selection criteria that can be used to estimate the number of regressor variables, none are used to estimate the number of interactions. Our simulations demonstrate that epistatic terms appearing in a model without the related main effects cause the standard model selection criteria to have a strong tendency to overestimate the number of interactions, and so the QTL number. With this as our motivation we investigate the behavior of the Schwarz Bayesian information criterion (BIC) by explaining the phenomenon of the overestimation and proposing a novel modification of BIC that allows the detection of main effects and pairwise interactions in a backcross population. Results of an extensive simulation study demonstrate that our modified version of BIC performs very well in practice. Our methodology can be extended to general populations and higher-order interactions.

Bayes Theorem↗

Variational learning and bits-back coding: an information-theoretic view to Bayesian learning.

The bits-back coding first introduced by Wallace in 1990 and later by Hinton and van Camp in 1993 provides an interesting link between Bayesian learning and information-theoretic minimum-description-length (MDL) learning approaches. The bits-back coding allows interpreting the cost function used in the variational Bayesian method called ensemble learning as a code length in addition to the Bayesian view of misfit of the posterior approximation and a lower bound of model evidence. Combining these two viewpoints provides interesting insights to the learning process and the functions of different parts of the model. In this paper, the problem of variational Bayesian learning of hierarchical latent variable models is used to demonstrate the benefits of the two views. The code-length interpretation provides new views to many parts of the problem such as model comparison and pruning and helps explain many phenomena occurring in learning.

Algorithms↗

A general class of Bayesian survival models with zero and nonzero cure fractions.

We propose a new class of survival models which naturally links a family of proper and improper population survival functions. The models resulting in improper survival functions are often referred to as cure rate models. This class of regression models is formulated through the Box-Cox transformation on the population hazard function and a proper density function. By adding an extra transformation parameter into the cure rate model, we are able to generate models with a zero cure rate, thus leading to a proper population survival function. A graphical illustration of the behavior and the influence of the transformation parameter on the regression model is provided. We consider a Bayesian approach which is motivated by the complexity of the model. Prior specification needs to accommodate parameter constraints due to the non-negativity of the survival function. Moreover, the likelihood function involves a complicated integral on the survival function, which may not have an analytical closed form, and thus makes the implementation of Gibbs sampling more difficult. We propose an efficient Markov chain Monte Carlo computational scheme based on Gaussian quadrature. The proposed method is illustrated with an example involving a melanoma clinical trial.

Adult↗

No free lunch for noise prediction.

No-free-lunch theorems have shown that learning algorithms cannot be universally good. We show that no free funch exists for noise prediction as well. We show that when the noise is additive and the prior over target functions is uniform, a prior on the noise distribution cannot be updated, in the Bayesian sense, from any finite data set. We emphasize the importance of a prior over the target function in order to justify superior performance for learning systems.

Algorithms↗

Bayesian nonparametric population models: formulation and comparison with likelihood approaches.

Population approaches to modeling pharmacokinetic and/or pharmacodynamic data attempt to separate the variability in observed data into within- and between-individual components. This is most naturally achieved via a multistage model. At the first stage of the model the data of a particular individual is modeled with each individual having his own set of parameters. At the second stage these individual parameters are assumed to have arisen from some unknown population distribution which we shall denote F. The importance of the choice of second stage distribution has led to a number of flexible approaches to the modeling of F. A nonparametric maximum likelihood estimate of F was suggested by Mallet whereas Davidian and Gallant proposed a semiparametric maximum likelihood approach where the maximum likelihood estimate is obtained over a smooth class of distributions. Previous Bayesian work has concentrated largely on F being assigned to a parametric family, typically the normal or Student's t. We describe a Bayesian nonparametric approach using the Dirichlet process. We use Markov chain Monte Carlo simulation to implement the procedure. We discuss each procedure and compare our approach with those of Mallet and Davidian and Gallant, using simulated data for a pharmacodynamic dose-response model.

Bayes Theorem↗

Advances in blind source separation (BSS) and independent component analysis (ICA) for nonlinear mixtures.

In this paper, we review recent advances in blind source separation (BSS) and independent component analysis (ICA) for nonlinear mixing models. After a general introduction to BSS and ICA, we discuss in more detail uniqueness and separability issues, presenting some new results. A fundamental difficulty in the nonlinear BSS problem and even more so in the nonlinear ICA problem is that they provide non-unique solutions without extra constraints, which are often implemented by using a suitable regularization. In this paper, we explore two possible approaches. The first one is based on structural constraints. Especially, post-nonlinear mixtures are an important special case, where a nonlinearity is applied to linear mixtures. For such mixtures, the ambiguities are essentially the same as for the linear ICA or BSS problems. The second approach uses Bayesian inference methods for estimating the best statistical parameters, under almost unconstrained models in which priors can be easily added. In the later part of this paper, various separation techniques proposed for post-nonlinear mixtures and general nonlinear mixtures are reviewed.

Bayes Theorem↗

Bayesian and maximum likelihood phylogenetic analyses of protein sequence data under relative branch-length differences and model violation.

BACKGROUND: Bayesian phylogenetic inference holds promise as an alternative to maximum likelihood, particularly for large molecular-sequence data sets. We have investigated the performance of Bayesian inference with empirical and simulated protein-sequence data under conditions of relative branch-length differences and model violation. RESULTS: With empirical protein-sequence data, Bayesian posterior probabilities provide more-generous estimates of subtree reliability than does the nonparametric bootstrap combined with maximum likelihood inference, reaching 100% posterior probability at bootstrap proportions around 80%. With simulated 7-taxon protein-sequence datasets, Bayesian posterior probabilities are somewhat more generous than bootstrap proportions, but do not saturate. Compared with likelihood, Bayesian phylogenetic inference can be as or more robust to relative branch-length differences for datasets of this size, particularly when among-sites rate variation is modeled using a gamma distribution. When the (known) correct model was used to infer trees, Bayesian inference recovered the (known) correct tree in 100% of instances in which one or two branches were up to 20-fold longer than the others. At ratios more extreme than 20-fold, topological accuracy of reconstruction degraded only slowly when only one branch was of relatively greater length, but more rapidly when there were two such branches. Under an incorrect model of sequence change, inaccurate trees were sometimes observed at less extreme branch-length ratios, and (particularly for trees with single long branches) such trees tended to be more inaccurate. The effect of model violation on accuracy of reconstruction for trees with two long branches was more variable, but gamma-corrected Bayesian inference nonetheless yielded more-accurate trees than did either maximum likelihood or uncorrected Bayesian inference across the range of conditions we examined. Assuming an exponential Bayesian prior on branch lengths did not improve, and under certain extreme conditions significantly diminished, performance. The two topology-comparison metrics we employed, edit distance and Robinson-Foulds symmetric distance, yielded different but highly complementary measures of performance. CONCLUSIONS: Our results demonstrate that Bayesian inference can be relatively robust against biologically reasonable levels of relative branch-length differences and model violation, and thus may provide a promising alternative to maximum likelihood for inference of phylogenetic trees from protein-sequence data.

Bayes Theorem↗

Use of Bayesian Markov Chain Monte Carlo methods to model cost-of-illness data.

It is well known that the modeling of cost data is often problematic due to the distribution of such data. Commonly observed problems include 1) a strongly right-skewed data distribution and 2) a significant percentage of zero-cost observations. This article demonstrates how a hurdle model can be implemented from a Bayesian perspective by means of Markov Chain Monte Carlo simulation methods using the freely available software WinBUGS. Assessment of model fit is addressed through the implementation of two cross-validation methods. The relative merits of this Bayesian approach compared to the classical equivalent are discussed in detail. To illustrate the methods described, patient-specific non-health-care resource-use data from a prospective longitudinal study and the Norfolk Arthritis Register (NOAR) are utilized for 218 individuals with early inflammatory polyarthritis (IP). The NOAR database also includes information on various patient-level covariates.

Arthritis↗

2D observers for human 3D object recognition?

In human object recognition, converging evidence has shown that subjects' performance depends on their familiarity with an object's appearance. The extent of such dependence is a function of the inter-object similarity. The more similar the objects are, the stronger this dependence will be and the more dominant the two-dimensional (2D) image-based information will be. However, the degree to which three-dimensional (3D) model-based information is used remains an area of strong debate. Previously the authors showed that all models with independent 2D templates that allowed 2D rotations in the image plane cannot account for human performance in discriminating novel object views. Here the authors derive an analytic formulation of a Bayesian model that gives rise to the best possible performance under 2D affine transformations and demonstrate that this model cannot account for human performance in 3D object discrimination. Relative to this model, human statistical efficiency is higher for novel views than for learned views, suggesting that human observers have used some 3D structural information.

Computer Simulation↗

Modeling compositional heterogeneity.

Compositional heterogeneity among lineages can compromise phylogenetic analyses, because models in common use assume compositionally homogeneous data. Models that can accommodate compositional heterogeneity with few extra parameters are described here, and used in two examples where the true tree is known with confidence. It is shown using likelihood ratio tests that adequate modeling of compositional heterogeneity can be achieved with few composition parameters, that the data may not need to be modelled with separate composition parameters for each branch in the tree. Tree searching and placement of composition vectors on the tree are done in a Bayesian framework using Markov chain Monte Carlo (MCMC) methods. Assessment of fit of the model to the data is made in both maximum likelihood (ML) and Bayesian frameworks. In an ML framework, overall model fit is assessed using the Goldman-Cox test, and the fit of the composition implied by a (possibly heterogeneous) model to the composition of the data is assessed using a novel tree-and model-based composition fit test. In a Bayesian framework, overall model fit and composition fit are assessed using posterior predictive simulation. It is shown that when composition is not accommodated, then the model does not fit, and incorrect trees are found; but when composition is accommodated, the model then fits, and the known correct phylogenies are obtained.

Base Composition↗

Bayesian difference refinement.

Interest in a pair of highly isomorphous structures often focuses on the differences between them. In cases where substantial correlated model errors exist or where there are differences in the quality of the two experimental data sets (cases quite common in macromolecular crystallography), independent refinement of the two structures does not lead to the most accurate estimate of the differences between them. An alternative procedure that has proven effective in some such cases is difference refinement, in which the residual between observed and calculated differences in structure-factor amplitudes between the two structures is minimized. A Bayesian approach has been used to extend the range of applicability of difference refinement to cases where there is only partial correlation in model errors and where the overlap between the data sets is limited. The resulting method, Bayesian difference refinement, uses residuals to be minimized that vary smoothly between difference refinement and independent refinement. When the errors in the two structural models are very similar, difference refinement is used; when they are very different, independent refinement is used; and when they are partially correlated, a combination of the two is used. The procedure is very simple to apply and does not significantly increase the computational demands of refinement.

Journal Article↗

Bayesian inference for prevalence in longitudinal two-phase studies.

We consider Bayesian inference and model selection for prevalence estimation using a longitudinal two-phase design in which subjects initially receive a low-cost screening test followed by an expensive diagnostic test conducted on several occasions. The change in the subject's diagnostic probability over time is described using four mixed-effects probit models in which the subject-specific effects are captured by latent variables. The computations are performed using Markov chain Monte Carlo methods. These models are then compared using the deviance information criterion. The methodology is illustrated with an analysis of alcohol and drug use in adolescents using data from the Great Smoky Mountains Study.

Adolescent↗

Bayesian approaches to meta-analysis of ROC curves.

A comparative review of important classic and Bayesian approaches to fixed-effects and random-effects meta-analysis of binormal ROC curves and areas underneath them is presented. The ROC analyses results of seven evaluation studies concerning the dexamethasone suppression test provide the basis for a worked example. Particular attention is given to fully Bayesian inference, a novelty in the ROC context, based on Gibbs samples from posterior distributions of hierarchical model parameters and related quantities. Fully Bayesian meta-analysis may properly account for the uncertainty associated with the model parameters, possibly incorporating prior knowledge and beliefs, and allows clinically intuitive predictions of unobserved study effects via calculation of posterior predictive densities. The effects of various different prior specifications (six noninformative as well as one informative) on the posterior estimates are investigated (sensitivity-analysis). Recommendations and suggestions for further research are made. Computer code for the more advanced methods may either be downloaded via the Internet or be found elsewhere.

Bayes Theorem↗

Application of a gamma model of absorption to oral cyclosporin.

BACKGROUND: Some drugs, such as cyclosporin, exhibit flat and delayed absorption profiles, with a correlation between the delay and the peak width. Such profiles can be described by an absorption model in which the absorption rate is derived from a gamma distribution (of which the classical first-order absorption model is a special case). OBJECTIVE: To develop a model for the pharmacokinetics of extravascular administration of cyclosporin and apply it to a study of the pharmacokinetics of cyclosporin microemulsion in stable renal transplant recipients. PATIENTS AND PARTICIPANTS: 21 renal transplant patients receiving oral cyclosporin microemulsion 75 to 175 mg twice daily. METHODS: The equation of the plasma concentration-time curve after oral administration was expressed as a convolution product between the absorption rate and a multi-exponential impulse response. The convolution integral was computed analytically and expressed in terms of the incomplete gamma function. Cyclosporin was assayed by liquid chromatography/mass spectrophotometry. The model was fitted by nonlinear regression, using a specially developed program. RESULTS: The gamma model yielded a good fit in all of the 21 patients studied. Attempts to fit the same data by a classical exponential with lag-time model failed in most patients. CONCLUSIONS: This model could simplify the Bayesian monitoring of cyclosporin therapy.

Area Under Curve↗

C. R. Henderson: contributions to predicting genetic merit.

The contributions of C. R. Henderson to the genetic evaluation of livestock have been widely accepted, utilized, and enhanced by animal breeders and statisticians. Not well known are the possible alternatives to BLUP that have been suggested by Henderson, such as biased predictors and Bayesian methodology to incorporate prior information. A search for rapid methods of inverting dominance and additive by additive genetic relationship matrices has taken place since Henderson first described his method for computing the inverse of the additive genetic relationship matrix. Accounting for inbreeding and selected base populations continues to be a problem. There is a need to derive accurate descriptions of selection processes and the appropriate selection model in order to solve the problems of cow culling, nonrandom association of sires with herd-year-seasons, and preferential treatment. Henderson solved many animal breeding problems and left hints for solving others, but there are still many difficult problems to be tackled that will have to be resolved without the availability of his generous advice.

Animals↗