Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Bayesian inferences in the Cox model for order-restricted hypotheses.

In studying the relationship between an ordered categorical predictor and an event time, it is standard practice to include dichotomous indicators of the different levels of the predictor in a Cox model. One can then use a multiple degree-of-freedom score or partial likelihood ratio test for hypothesis testing. Often, interest focuses on comparing the null hypothesis of no difference to an order-restricted alternative, such as a monotone increase across levels of a predictor. This article proposes a Bayesian approach for addressing hypotheses of this type. We reparameterize the Cox model in terms of a cumulative product of parameters having conjugate prior densities, consisting of mixtures of point masses at one, and truncated gamma densities. Due to the structure of the model, posterior computation can proceed via a simple and efficient Gibbs sampling algorithm. Posterior probabilities for the global null hypothesis and subhypotheses, comparing the hazards for specific groups, can be calculated directly from the output of a single Gibbs chain. The approach allows for level sets across which a predictor has no effect. Generalizations to multiple predictors are described, and the method is applied to a study of emergency medical treatment for stroke.

Bayes Theorem↗

Speech utterance clustering based on the maximization of within-cluster homogeneity of speaker voice characteristics.

This paper investigates the problem of how to partition unknown speech utterances into a set of clusters, such that each cluster consists of utterances from only one speaker, and the number of clusters reflects the unknown speaker population size. The proposed method begins by specifying a certain number of clusters, corresponding to one of the possible speaker population sizes, and then maximizes the level of overall within-cluster homogeneity of the speakers' voice characteristics. The within-cluster homogeneity is characterized by the likelihood probability that a cluster model, trained using all the utterances within a cluster, matches each of the within-cluster utterances. To attain the maximal sum of likelihood probabilities for all utterances, the proposed method applies a genetic algorithm to determine the cluster in which each utterance should be located. For greater computational efficiency, also proposed is a clustering criterion that approximates the likelihood probability with a divergence-based model similarity between a cluster and each of the within-cluster utterances. The clustering method then examines various legitimate numbers of clusters by adapting the Bayesian information criterion to determine the most likely speaker population size. The experimental results show the superiority of the proposed method over conventional methods based on hierarchical clustering.

Algorithms↗

Assessing heterogeneity and correlation of paired failure times with the bivariate frailty model.

We consider bivariate survival times for heterogeneous populations, where heterogeneity induces deviations in an individual's risk of an event as well as associations between survival times. The heterogeneity is characterized by a bivariate frailty model. We measure the heterogeneity effects through deviations associated with hazard functions and an association function defined through the conditional hazard functions: the cross-ratio function proposed by Oakes. We show how the deviation and association measures are determined by the frailty distribution. A Gibbs sampling method is developed for Bayesian inferences on regression coefficients, frailty parameters and the heterogeneity measures. The method is applied to a mental health care data set.

Algorithms↗

Statistical analysis of pharmacokinetic models in dynamic contrast-enhanced magnetic resonance imaging.

This paper assesses the estimation of kinetic parameters from dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI). Asymptotic results from likelihood-based nonlinear regression are compared with results derived from the posterior distribution using Bayesian estimation, along with the output from an established software package (MRIW). By using the estimated error from kinetic parameters, it is possible to produce more accurate clinical statistics, such as tumor size, for patients with breast tumors. Further analysis has also shown that Bayesian methods are more accurate and do not suffer from convergence problems, but at a higher computational cost.

Breast Neoplasms↗

Bias and precision in QST estimates: problems and some solutions.

Comparison of population differentiation in neutral marker genes and in genes coding quantitative traits by means of F(ST) and Q(ST) indexes has become commonplace practice. While the properties and estimation of F(ST) have been the subject of much interest, little is known about the precision and possible bias in Q(ST) estimates. Using both simulated and real data, we investigated the precision and bias in Q(ST) estimates and various methods of estimating the precision. We found that precision of Q(ST) estimates for typical data sets (i.e., with <20 populations) was poor. Of the methods for estimating the precision, a simulation method, a parametric bootstrap, and the Bayesian approach returned the most precise estimates of the confidence intervals.

Animals↗

A calculator program for clinical application of the Bayesian method of predicting plasma drug levels.

A pharmacokinetic program that allows individualization of drug dosage regimens through the Bayesian method is described. The program, which is designed for the Hewlett-Packard HP-41 CV calculator, is based upon the one-compartment open model with either instantaneous or zero-order absorption. Individualized estimation of the patient's kinetic parameters (clearance and volume of distribution) is performed by analyzing the plasma levels measured in the patient as well as considering the population data of the drug. After estimating the individual kinetic parameters by the Bayesian method, the program predicts the dosage regimen that will elicit the desired peak and trough plasma levels at steady state. For comparison purposes, the least-squares estimates for clearance and volume of distribution are calculated, and dosage prediction can also be made on the basis of the least-squares estimates. The least-squares estimates can be used to calculate population pharmacokinetic parameters according to the Standard Two-Stage method. Several examples of clinical use of the program are presented. The examples refer to patients with classic hemophilia who were treated with Factor VIII concentrates. In these patients, the Bayesian kinetic parameters of Factor VIII have been estimated through the calculator program. The Bayesian parameter estimates generated by the HP-41 have been compared with those determined by a Bayesian program (ADVISE) designed for microcomputers.

Computers↗

Differential and trajectory methods for time course gene expression data.

MOTIVATION: The issue of high dimensionality in microarray data has been, and remains, a hot topic in statistical and computational analysis. Efficient gene filtering and differentiation approaches can reduce the dimensions of data, help to remove redundant genes and noises, and highlight the most relevant genes that are major players in the development of certain diseases or the effect of drug treatment. The purpose of this study is to investigate the efficiency of parametric (including Bayesian and non-Bayesian, linear and non-linear), non-parametric and semi-parametric gene filtering methods through the application of time course microarray data from multiple sclerosis patients being treated with interferon-beta-1a. The analysis of variance with bootstrapping (parametric), class dispersion (semi-parametric) and Pareto (non-parametric) with permutation methods are presented and compared for filtering and finding differentially expressed genes. The Bayesian linear correlated model, the Bayesian non-linear model the and non-Bayesian mixed effects model with bootstrap were also developed to characterize the differential expression patterns. Furthermore, trajectory-clustering approaches were developed in order to investigate the dynamic patterns and inter-dependency of drug treatment effects on gene expression. RESULTS: Results show that the presented methods performed significant differently but all were adequate in capturing a small number of the potentially relevant genes to the disease. The parametric method, such as the mixed model and two Bayesian approaches proved to be more conservative. This may because these methods are based on overall variation in expression across all time points. The semi-parametric (class dispersion) and non-parametric (Pareto) methods were appropriate in capturing variation in expression from time point to time point, thereby making them more suitable for investigating significant monotonic changes and trajectories of changes in gene expressions in time course microarray data. Also, the non-linear Bayesian model proved to be less conservative than linear Bayesian correlated growth models to filter out the redundant genes, although the linear model showed better fit than non-linear model (smaller DIC). We also report the trajectories of significant genes-since we have been able to isolate trajectories of genes whose regulations appear to be inter-dependent.

Computer Simulation↗

A Bayesian network model for protein fold and remote homologue recognition.

MOTIVATION: The Bayesian network approach is a framework which combines graphical representation and probability theory, which includes, as a special case, hidden Markov models. Hidden Markov models trained on amino acid sequence or secondary structure data alone have been shown to have potential for addressing the problem of protein fold and superfamily classification. RESULTS: This paper describes a novel implementation of a Bayesian network which simultaneously learns amino acid sequence, secondary structure and residue accessibility for proteins of known three-dimensional structure. An awareness of the errors inherent in predicted secondary structure may be incorporated into the model by means of a confusion matrix. Training and validation data have been derived for a number of protein superfamilies from the Structural Classification of Proteins (SCOP) database. Cross validation results using posterior probability classification demonstrate that the Bayesian network performs better in classifying proteins of known structural superfamily than a hidden Markov model trained on amino acid sequences alone.

Amino Acid Sequence↗

A hierarchical Bayesian model to predict the duration of immunity to Haemophilus influenzae type b.

A hierarchical Bayesian regression model is fitted to longitudinal data on Haemophilus influenzae type b (Hib) serum antibodies. To estimate the decline rate of the antibody concentration, the model accommodates the possibility of unobserved subclinical infections with Hib bacteria that cause increasing concentrations during the study period. The computations rely on Markov chain Monte Carlo simulation of the joint posterior distribution of the model parameters. The model is used to predict the duration of immunity to subclinical Hib infection and to a serious invasive Hib disease.

Antibodies, Bacterial↗

Numerical evaluation of cytologic data. V. Bivariate distributions and the Bayesian decision Boundary.

The evaluation of cytologic data often involves the classification of observations into alternative categories (data sets). Plotting the elliptical contours of bivariate distributions provides immediate insight into the structure of data sets and their mutual relations. In this paper, the computation of tolerance ellipses and confidence ellipses for bivariate distribution is demonstrated, and the finding of a Bayesian decision boundary between two bivariate distributions is illustrated.

Bayes Theorem↗

Hypothesis testing and Bayesian estimation using a sigmoid Emax model applied to sparse dose-response designs.

Application of a sigmoid Emax model is described for the assessment of dose-response with designs containing a small number of doses (typically, three to six). The expanded model is a common Emax model with a power (Hill) parameter applied to dose and the ED50 parameter. The model will be evaluated following a strategy proposed by Bretz et al. (2005). The sigmoid Emax model is used to create several contrasts that have high power to detect an increasing trend from placebo. Alpha level for the hypothesis of no dose-response is controlled using multiple comparison methods applied to the p-values obtained from the contrasts. Subsequent to establishing drug activity, Bayesian methods are used to estimate the dose-response curve from the sparse dosing design. Bayesian estimation applied to the sigmoid model represents uncertainty in model selection that is missed when a single simpler model is selected from a collection of non-nested models. The goal is to base model selection on substantive knowledge and broad experience with dose-response relationships rather than criteria selected to ensure convergence of estimators. Bayesian estimation also addresses deficiencies in confidence intervals and tests derived from asymptotic-based maximum likelihood estimation when some parameters are poorly determined, which is typical for data from common dose-response designs.

Algorithms↗

A population approach to initial dose selection.

Before a drug can be marketed, an initial dose must be established. Sheiner et al. argue that a population approach leads to the most informed and rational decision making. We discuss the choice of an initial dose from both a predictive and estimative viewpoint. Our criteria are based upon evaluating the probabilities that a patient from the specified population obtains a response that is at least of a specified size. We demonstrate the approach using a simulation study and compare estimation of population parameters and initial dose using Bayesian and likelihood-based methods.

Bayes Theorem↗

GenSo-FDSS: a neural-fuzzy decision support system for pediatric ALL cancer subtype identification using gene expression data.

OBJECTIVE: Acute lymphoblastic leukemia (ALL) is the most common malignancy of childhood, representing nearly one third of all pediatric cancers. Currently, the treatment of pediatric ALL is centered on tailoring the intensity of the therapy applied to a patient's risk of relapse, which is linked to the type of leukemia the patient has. Hence, accurate and correct diagnosis of the various leukemia subtypes becomes an important first step in the treatment process. Recently, gene expression profiling using DNA microarrays has been shown to be a viable and accurate diagnostic tool to identify the known prognostically important ALL subtypes. Thus, there is currently a huge interest in developing autonomous classification systems for cancer diagnosis using gene expression data. This is to achieve an unbiased analysis of the data and also partly to handle the large amount of genetic information extracted from the DNA microarrays. METHODOLOGY: Generally, existing medical decision support systems (DSS) for cancer classification and diagnosis are based on traditional statistical methods such as Bayesian decision theory and machine learning models such as neural networks (NN) and support vector machine (SVM). Though high accuracies have been reported for these systems, they fall short on certain critical areas. These included (a) being able to present the extracted knowledge and explain the computed solutions to the users; (b) having a logical deduction process that is similar and intuitive to the human reasoning process; and (c) flexible enough to incorporate new knowledge without running the risk of eroding old but valid information. On the other hand, a neural fuzzy system, which is synthesized to emulate the human ability to learn and reason in the presence of imprecise and incomplete information, has the ability to overcome the above-mentioned shortcomings. However, existing neural fuzzy systems have their own limitations when used in the design and implementation of DSS. Hence, this paper proposed the use of a novel neural fuzzy system: the generic self-organising fuzzy neural network (GenSoFNN) with truth-value restriction (TVR) fuzzy inference, as a fuzzy DSS (denoted as GenSo-FDSS) for the classification of ALL subtypes using gene expression data. RESULTS AND CONCLUSION: The performance of the GenSo-FDSS system is encouraging when benchmarked against those of NN, SVM and the K-nearest neighbor (K-NN) classifier. On average, a classification rate of above 90% has been achieved using the GenSo-FDSS system.

Algorithms↗

Bayesian networks for knowledge discovery in large datasets: basics for nurse researchers.

The growth of nursing databases necessitates new approaches to data analyses. These databases, which are known to be massive and multidimensional, easily exceed the capabilities of both human cognition and traditional analytical approaches. One innovative approach, knowledge discovery in large databases (KDD), allows investigators to analyze very large data sets more comprehensively in an automatic or a semi-automatic manner. Among KDD techniques, Bayesian networks, a state-of-the art representation of probabilistic knowledge by a graphical diagram, has emerged in recent years as essential for pattern recognition and classification in the healthcare field. Unlike some data mining techniques, Bayesian networks allow investigators to combine domain knowledge with statistical data, enabling nurse researchers to incorporate clinical and theoretical knowledge into the process of knowledge discovery in large datasets. This tailored discussion presents the basic concepts of Bayesian networks and their use as knowledge discovery tools for nurse researchers.

Artificial Intelligence↗

Algebraic geometrical methods for hierarchical learning machines.

Hierarchical learning machines such as layered perceptrons, radial basis functions, Gaussian mixtures are non-identifiable learning machines, whose Fisher information matrices are not positive definite. This fact shows that conventional statistical asymptotic theory cannot be applied to neural network learning theory, for example either the Bayesian a posteriori probability distribution does not converge to the Gaussian distribution, or the generalization error is not in proportion to the number of parameters. The purpose of this paper is to overcome this problem and to clarify the relation between the learning curve of a hierarchical learning machine and the algebraic geometrical structure of the parameter space. We establish an algorithm to calculate the Bayesian stochastic complexity based on blowing-up technology in algebraic geometry and prove that the Bayesian generalization error of a hierarchical learning machine is smaller than that of a regular statistical model, even if the true distribution is not contained in the parametric model.

Algorithms↗

Model-free analysis of protein dynamics: assessment of accuracy and model selection protocols based on molecular dynamics simulation.

The popular model-free approach to analyze NMR relaxation measurements has been examined using artificial amide (15)N relaxation data sets generated from a 10 nanosecond molecular dynamics trajectory of a dihydrofolate reductase ternary complex in explicit water. With access to a detailed picture of the underlying internal motions, the efficacy of model-free analysis and impact of model selection protocols on the interpretation of NMR data can be studied. In the limit of uncorrelated global tumbling and internal motions, fitting the relaxation data to the model-free models can recover a significant amount of quantitative information on the internal dynamics. Despite a slight overestimation, the generalized order parameter is quite accurately determined. However, the model-free analysis appears to be insensitive to the presence of nanosecond time scale motions with relatively small magnitude. For such cases, the effective correlation time can be significantly underestimated. As a result, proteins appear to be more rigid than they really are. The model selection protocols have a major impact on the information one can reliably obtain. The commonly employed protocol based on step-up hypothesis testing has severe drawbacks of oversimplification and underfitting. The consequences are that the order parameter is more severely overestimated and the correlation time more severely underestimated. Instead, model selection based on Bayesian Information Criteria (BIC), recently introduced to the model-free analysis by d'Auvergne and Gooley (2003), provides a better balance between bias and variance. More appropriate models can be selected, leading to improved estimate of both the order parameter and correlation time. In addition, the computational cost is significantly reduced and subjective parameters such as the significance level are unnecessary.

Anisotropy↗

Bayesian color constancy.

The problem of color constancy may be solved if we can recover the physical properties of illuminants and surfaces from photosensor responses. We consider this problem within the framework of Bayesian decision theory. First, we model the relation among illuminants, surfaces, and photosensor responses. Second, we construct prior distributions that describe the probability that particular illuminants and surfaces exist in the world. Given a set of photosensor responses, we can then use Bayes's rule to compute the posterior distribution for the illuminants and the surfaces in the scene. There are two widely used methods for obtaining a single best estimate from a posterior distribution. These are maximum a posteriori (MAP) and minimum mean-square-error (MMSE) estimation. We argue that neither is appropriate for perception problems. We describe a new estimator, which we call the maximum local mass (MLM) estimate, that integrates local probability density. The new method uses an optimality criterion that is appropriate for perception tasks: It finds the most probable approximately correct answer. For the case of low observation noise, we provide an efficient approximation. We develop the MLM estimator for the color-constancy problem in which flat matte surfaces are uniformly illuminated. In simulations we show that the MLM method performs better than the MAP estimator and better than a number of standard color-constancy algorithms. We note conditions under which even the optimal estimator produces poor estimates: when the spectral properties of the surfaces in the scene are biased.

Color Perception↗