Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Model evaluation and spatial interpolation by Bayesian combination of observations with outputs from numerical models.

Constructing maps of dry deposition pollution levels is vital for air quality management, and presents statistical problems typical of many environmental and spatial applications. Ideally, such maps would be based on a dense network of monitoring stations, but this does not exist. Instead, there are two main sources of information for dry deposition levels in the United States: one is pollution measurements at a sparse set of about 50 monitoring stations called CASTNet, and the other is the output of the regional scale air quality models, called Models-3. A related problem is the evaluation of these numerical models for air quality applications, which is crucial for control strategy selection. We develop formal methods for combining sources of information with different spatial resolutions and for the evaluation of numerical models. We specify a simple model for both the Models-3 output and the CASTNet observations in terms of the unobserved ground truth, and we estimate the model in a Bayesian way. This provides improved spatial prediction via the posterior distribution of the ground truth, allows us to validate Models-3 via the posterior predictive distribution of the CASTNet observations, and enables us to remove the bias in the Models-3 output. We apply our methods to data on SO2 concentrations, and we obtain high-resolution SO2 distributions by combining observed data with model output. We also conclude that the numerical models perform worse in areas closer to power plants, where the SO2 values are overestimated by the models.

Air Pollution↗

Bayesian Gaussian process classification with the EM-EP algorithm.

Gaussian process classifiers (GPCs) are Bayesian probabilistic kernel classifiers. In GPCs, the probability of belonging to a certain class at an input location is monotonically related to the value of some latent function at that location. Starting from a Gaussian process prior over this latent function, data are used to infer both the posterior over the latent function and the values of hyperparameters to determine various aspects of the function. Recently, the expectation propagation (EP) approach has been proposed to infer the posterior over the latent function. Based on this work, we present an approximate EM algorithm, the EM-EP algorithm, to learn both the latent function and the hyperparameters. This algorithm is found to converge in practice and provides an efficient Bayesian framework for learning hyperparameters of the kernel. A multiclass extension of the EM-EP algorithm for GPCs is also derived. In the experimental results, the EM-EP algorithms are as good or better than other methods for GPCs or Support Vector Machines (SVMs) with cross-validation.

Algorithms↗

Active and dynamic information fusion for multisensor systems with dynamic Bayesian networks.

Many information fusion applications are often characterized by a high degree of complexity because: (1) data are often acquired from sensors of different modalities and with different degrees of uncertainty; (2) decisions must be made efficiently; and (3) the world situation evolves over time. To address these issues, we propose an information fusion framework based on dynamic Bayesian networks to provide active, dynamic, purposive and sufficing information fusion in order to arrive at a reliable conclusion with reasonable time and limited resources. The proposed framework is suited to applications where the decision must be made efficiently from dynamically available information of diverse and disparate sources.

Algorithms↗

Prognostic factors for non-cleaved follicular center-cell lymphomas and immunoblastic sarcoma. A Bayesian approach.

The Bayesian multivariate statistical method was applied to determine the relative strength and optimal combination of 17 variables in predicting survival of 151 patients with non-Hodgkin's lymphomas assigned as non-cleaved follicular center-cell and immunoblastic sarcoma types according to the classification of Lukes & Collins. Considering all the factors simultaneously, the analysis showed that the combination of stage, Hb level and location of the lymphoma was included in the best predictive model at each survival time studied. Additional factors were erythrocyte sedimentation rate, thrombocyte count and leucocyte count. Of the histological variables, only growth pattern and mitotic ratio in the biopsy specimen remained significant. At manually controlled computer simulation with these best indicators, this model would have given a correct classification for 69-78% of the patients at the 4 survival times studied. One can thus expect about 70% correct prognoses using this model.

Adult↗

Combining microarrays and biological knowledge for estimating gene networks via bayesian networks.

We propose a statistical method for estimating a gene network based on Bayesian networks from microarray gene expression data together with biological knowledge including protein-protein interactions, protein-DNA interactions, binding site information, existing literature and so on. Microarray data do not contain enough information for constructing gene networks accurately in many cases. Our method adds biological knowledge to the estimation method of gene networks under a Bayesian statistical framework, and also controls the trade-off between microarray information and biological knowledge automatically. We conduct Monte Carlo simulations to show the effectiveness of the proposed method. We analyze Saccharomyces cerevisiae gene expression data as an application.

Algorithms↗

Detecting recombination in evolving nucleotide sequences.

BACKGROUND: Genetic recombination can produce heterogeneous phylogenetic histories within a set of homologous genes. These recombination events can be obscured by subsequent residue substitutions, which consequently complicate their detection. While there are many algorithms for the identification of recombination events, little is known about the effects of subsequent substitutions on the accuracy of available recombination-detection approaches. RESULTS: We assessed the effect of subsequent substitutions on the detection of simulated recombination events within sets of four nucleotide sequences under a homogeneous evolutionary model. The amount of subsequent substitutions per site, prior evolutionary history of the sequences, and reciprocality or non-reciprocality of the recombination event all affected the accuracy of the recombination-detecting programs examined. Bayesian phylogenetic-based approaches showed high accuracy in detecting evidence of recombination event and in identifying recombination breakpoints. These approaches were less sensitive to parameter settings than other methods we tested, making them easier to apply to various data sets in a consistent manner. CONCLUSION: Post-recombination substitutions tend to diminish the predictive accuracy of recombination-detecting programs. The best method for detecting recombined regions is not necessarily the most accurate in identifying recombination breakpoints. For difficult detection problems involving highly divergent sequences or large data sets, different types of approach can be run in succession to increase efficiency, and can potentially yield better predictive accuracy than any single method used in isolation.

Algorithms↗

Flexible random-effects models using Bayesian semi-parametric models: applications to institutional comparisons.

Random effects models are used in many applications in medical statistics, including meta-analysis, cluster randomized trials and comparisons of health care providers. This paper provides a tutorial on the practical implementation of a flexible random effects model based on methodology developed in Bayesian non-parametrics literature, and implemented in freely available software. The approach is applied to the problem of hospital comparisons using routine performance data, and among other benefits provides a diagnostic to detect clusters of providers with unusual results, thus avoiding problems caused by masking in traditional parametric approaches. By providing code for Winbugs we hope that the model can be used by applied statisticians working in a wide variety of applications.

Bayes Theorem↗

Site-specific evolutionary rate inference: taking phylogenetic uncertainty into account.

The evolutionary rate at an amino acid site is indicative of how conserved this site is and, in turn, allows evaluating the importance of this site in maintaining the structure/function of the protein. When evolutionary rates are estimated, one must reconstruct the phylogenetic tree describing the evolutionary relationship among the sequences under study. However, if the inferred phylogenetic tree is incorrect, it can lead to erroneous site-specific rate estimates. Here we describe a novel Bayesian method that uses Markov chain Monte Carlo methodology to integrate over the space of all possible trees and model parameters. By doing so, the method considers alternative evolutionary scenarios weighted by their posterior probabilities. We show that this comprehensive evolutionary approach is superior over methods that are based on only a single tree. We illustrate the potential of our algorithm by analyzing the conservation pattern of the potassium channel protein family.

Algorithms↗

Predicting the daily prothrombin time response to warfarin.

Our objective was to evaluate the effectiveness of a computer program to predict daily prothrombin time (PT) response to warfarin therapy using prospectively collected data. The program's predictive performance (precision) and accuracy (bias) were evaluated using fraction mean absolute error and fraction mean error, respectively. We analyzed data from 40 patients using from zero to nine PT feedbacks. The fraction mean absolute error varied from 0.058 to 0.13. The program utilized a pharmacokinetic/pharmacodynamic Bayesian forecasting system to predict prothrombin response.

Aged↗

Bayesian extensions of the Tobit model for analyzing measures of health status.

Self-reported health status is often measured using utility indices that provide a score intended to summarize an individual's health. Measurements of health status can be subject to a ceiling effect. Frequently, researchers want to examine relationships between determinants of health and measures of health status. In this article, Bayesian extensions of the classical Tobit model are used to study the relationship between health status and predictors of health. The author examined models where the conditional distribution of health status was either normal or lognormal, and allowed for both homoscedasticity and heteroscedasticity. Bayes factors were then used to compare the evidence for a given model against that for a competing model. The author found very strong evidence that the distribution of the Health Utilities Index, conditional on age, gender, income adequacy, and number of chronic conditions, was normal with nonuniform variance, compared to the competing models.

Bayes Theorem↗

Bayesian shrinkage estimation of quantitative trait loci parameters.

Mapping multiple QTL is a typical problem of variable selection in an oversaturated model because the potential number of QTL can be substantially larger than the sample size. Currently, model selection is still the most effective approach to mapping multiple QTL, although further research is needed. An alternative approach to analyzing an oversaturated model is the shrinkage estimation in which all candidate variables are included in the model but their estimated effects are forced to shrink toward zero. In contrast to the usual shrinkage estimation where all model effects are shrunk by the same factor, we develop a Bayesian method that allows the shrinkage factor to vary across different effects. The new shrinkage method forces marker intervals that contain no QTL to have estimated effects close to zero whereas intervals containing notable QTL have estimated effects subject to virtually no shrinkage. We demonstrate the method using both simulated and real data for QTL mapping. A simulation experiment with 500 backcross (BC) individuals showed that the method can localize closely linked QTL and QTL with effects as small as 1% of the phenotypic variance of the trait. The method was also used to map QTL responsible for wound healing in a family of a (MRL/MPJ x SJL/J) cross with 633 F(2) mice derived from two inbred lines.

Animals↗

Late potential recognition by artificial neural networks.

Ventricular late potentials (LP's) are high-frequency low-amplitude signals obtained from signal-averaged electrocardiograms (ECG's) [SAECG's]. LP's are useful in identifying patients prone to ventricular tachycardia (VT), spontaneous or inducible during electrophysiology testing. A combination of self-organizing and supervised artificial neural network (ANN) models was developed to identify patients with a positive electrophysiology (PEP) test for inducible ventricular tachycardia from patients with a negative electrophysiology (NEP) test using LP's. We have added morphology information of vector magnitude waveform to original set of three time-domain features of LP's, which are total QRS duration (TQRSD), high-frequency low-amplitude signal duration (HFLAD), and root-mean-square voltage (RMSV). Pattern recognition results from an ANN model with this combination feature set are superior to the results from Bayesian classification model based on conventional three time-domain features of SAECG. In order to increase the robustness of the recognition, a filtered QRS offset point is randomly shifted +/- 8 ms to form a fuzzy training set, which was to simulate the possible error in detecting QRS offset point of filtered SAECG. We also found that nonlinear transformation through the hidden layer of developed ANN model could increase Euclidean distance between PEP and NEP patterns.

Algorithms↗

Evidence for multiple reversals of asymmetric mutational constraints during the evolution of the mitochondrial genome of metazoa, and consequences for phylogenetic inferences.

Mitochondrial DNA (mtDNA) sequences are comonly used for inferring phylogenetic relationships. However, the strand-specific bias in the nucleotide composition of the mtDNA, which is thought to reflect assymetric mutational constraints, combined with the important compositional heterogeneity among taxa, are known to be highly problematic for phylogenetic analyses. Here, nucleotide composition was compared across 49 species of Metazoa (34 arthropods, 2 annelids, 2 molluscs, and 11 deuterosomes), and analyzed for a mtDNA fragment including six protein-coding genes, i.e., atp6, atp8, cox1, cox2, cox3, and nad2. The analyses show that most metazoan species present a clear strand assymetry, where one strand is biased in favor of A and C, whereas the other strand has reverse bias, i.e. in favor of T and G. the origin of this strand bias can be related to assymetric mutational constraints involving deaminations of A and C nucleotides during the replication and/or transcription processes. The analyses reveal that six unrelated genera are characterized by a reversal of the usual strand bias, i.e., Argiope (Araneae), Euscorpius (Scorpiones), Tigrioupus (Maxillopoda), Branchiostoma (Cephalochordata) Florometra (Echinodermata), and Katharina (Mollusca). It is proposed that assymetric mutational constraints have been independantly reversed in these six genera, through an inversion of the control region, i.e., the region that contains most regulatory elements for replication and transcription of the mtDNA. We show that reversals of assymetric mutational constraints have dramatic consequences on the phylogenetic analyses, as taxa characterized by reverse strand bias tend to group together due to long-branch attraction artifacts. We propose a new method for limiting this specific problem in tree reconstruction under the Bayesian approach. We apply our method to deal with the question of phylogenetic relationships of the major lineages of Arthropoda, This new approach provides a better congruence with nuclear analyses based on mtDNA sequences, our data suggest that Chelicerata, Crustacea, Myriapoda, Pancrustacea, and Paradoxopoda are monophyletic.

Animals↗

Classical and Bayesian inference in neuroimaging: theory.

This paper reviews hierarchical observation models, used in functional neuroimaging, in a Bayesian light. It emphasizes the common ground shared by classical and Bayesian methods to show that conventional analyses of neuroimaging data can be usefully extended within an empirical Bayesian framework. In particular we formulate the procedures used in conventional data analysis in terms of hierarchical linear models and establish a connection between classical inference and parametric empirical Bayes (PEB) through covariance component estimation. This estimation is based on an expectation maximization or EM algorithm. The key point is that hierarchical models not only provide for appropriate inference at the highest level but that one can revisit lower levels suitably equipped to make Bayesian inferences. Bayesian inferences eschew many of the difficulties encountered with classical inference and characterize brain responses in a way that is more directly predicated on what one is interested in. The motivation for Bayesian approaches is reviewed and the theoretical background is presented in a way that relates to conventional methods, in particular restricted maximum likelihood (ReML). This paper is a technical and theoretical prelude to subsequent papers that deal with applications of the theory to a range of important issues in neuroimaging. These issues include; (i) Estimating nonsphericity or variance components in fMRI time-series that can arise from serial correlations within subject, or are induced by multisubject (i.e., hierarchical) studies. (ii) Spatiotemporal Bayesian models for imaging data, in which voxels-specific effects are constrained by responses in other voxels. (iii) Bayesian estimation of nonlinear models of hemodynamic responses and (iv) principled ways of mixing structural and functional priors in EEG source reconstruction. Although diverse, all these estimation problems are accommodated by the PEB framework described in this paper.

Algorithms↗

Bayesian model search for mixture models based on optimizing variational bounds.

When learning a mixture model, we suffer from the local optima and model structure determination problems. In this paper, we present a method for simultaneously solving these problems based on the variational Bayesian (VB) framework. First, in the VB framework, we derive an objective function that can simultaneously optimize both model parameter distributions and model structure. Next, focusing on mixture models, we present a deterministic algorithm to approximately optimize the objective function by using the idea of the split and merge operations which we previously proposed within the maximum likelihood framework. Then, we apply the method to mixture of expers (MoE) models to experimentally show that the proposed method can find the optimal number of experts of a MoE while avoiding local maxima.

Algorithms↗

Broad-based quantitative structure-activity relationship modeling of potency and selectivity of farnesyltransferase inhibitors using a Bayesian regularized neural network.

Inhibitors of the enzyme farnesyltransferase show potential as novel anticancer agents. There are many known inhibitors, but efforts to build predictive SAR models have been hampered by the structural diversity and flexibility of inhibitors. We have undertaken for the first time a QSAR study of the potency and selectivity of a large, diverse data set of farnesyltransferase inhibitors. We used novel molecular descriptors based on binned atomic properties and invariants of molecular matrices and a robust, nonlinear QSAR mapping paradigm, the Bayesian regularized neural network. We have built robust QSAR models of farnesyltransferase inhibition, geranylgeranyltransferase inhibition, and in vivo data. We have derived a novel selectivity index that allows us to model potency and selectivity simultaneously and have built robust QSAR models using this index that have the potential to discover new potent and selective inhibitors.

Alkyl and Aryl Transferases↗

An accurate and efficient bayesian method for automatic segmentation of brain MRI.

Automatic three-dimensional (3-D) segmentation of the brain from magnetic resonance (MR) scans is a challenging problem that has received an enormous amount of attention lately. Of the techniques reported in the literature, very few are fully automatic. In this paper, we present an efficient and accurate, fully automatic 3-D segmentation procedure for brain MR scans. It has several salient features; namely, the following. 1) Instead of a single multiplicative bias field that affects all tissue intensities, separate parametric smooth models are used for the intensity of each class. 2) A brain atlas is used in conjunction with a robust registration procedure to find a nonrigid transformation that maps the standard brain to the specimen to be segmented. This transformation is then used to: segment the brain from nonbrain tissue; compute prior probabilities for each class at each voxel location and find an appropriate automatic initialization. 3) Finally, a novel algorithm is presented which is a variant of the expectation-maximization procedure, that incorporates a fast and accurate way to find optimal segmentations, given the intensity models along with the spatial coherence assumption. Experimental results with both synthetic and real data are included, as well as comparisons of the performance of our algorithm with that of other published methods.

Algorithms↗

A Bayesian approach to joint feature selection and classifier design.

This paper adopts a Bayesian approach to simultaneously learn both an optimal nonlinear classifier and a subset of predictor variables (or features) that are most relevant to the classification task. The approach uses heavy-tailed priors to promote sparsity in the utilization of both basis functions and features; these priors act as regularizers for the likelihood function that rewards good classification on the training data. We derive an expectation-maximization (EM) algorithm to efficiently compute a maximum a posteriori (MAP) point estimate of the various parameters. The algorithm is an extension of recent state-of-the-art sparse Bayesian classifiers, which in turn can be seen as Bayesian counterparts of support vector machines. Experimental comparisons using kernel classifiers demonstrate both parsimonious feature selection and excellent classification accuracy on a range of synthetic and benchmark data sets.

Algorithms↗