Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Context-specific Bayesian clustering for gene expression data.

The recent growth in genomic data and measurements of genome-wide expression patterns allows us to apply computational tools to examine gene regulation by transcription factors. In this work, we present a class of mathematical models that help in understanding the connections between transcription factors and functional classes of genes based on genetic and genomic data. Such a model represents the joint distribution of transcription factor binding sites and of expression levels of a gene in a unified probabilistic model. Learning a combined probability model of binding sites and expression patterns enables us to improve the clustering of the genes based on the discovery of putative binding sites and to detect which binding sites and experiments best characterize a cluster. To learn such models from data, we introduce a new search method that rapidly learns a model according to a Bayesian score. We evaluate our method on synthetic data as well as on real life data and analyze the biological insights it provides. Finally, we demonstrate the applicability of the method to other data analysis problems in gene expression data.

Bayes Theorem↗

Eutrophication risk assessment in coastal embayments using simple statistical models.

A statistical methodology is proposed for assessing the risk of eutrophication in marine coastal embayments. The procedure followed was the development of regression models relating the levels of chlorophyll a (Chl) with the concentration of the limiting nutrient--usually nitrogen--and the renewal rate of the systems. The method was applied in the Gulf of Gera, Island of Lesvos, Aegean Sea and a surrogate for renewal rate was created using the Canberra metric as a measure of the resemblance between the Gulf and the oligotrophic waters of the open sea in terms of their physical, chemical and biological properties. The Chl-total dissolved nitrogen-renewal rate regression model was the most significant, accounting for 60% of the variation observed in Chl. Predicted distributions of Chl for various combinations of the independent variables, based on Bayesian analysis of the models, enabled comparison of the outcomes of specific scenarios of interest as well as further analysis of the system dynamics. The present statistical approach can be used as a methodological tool for testing the resilience of coastal ecosystems under alternative managerial schemes and levels of exogenous nutrient loading.

Bayes Theorem↗

Modeling the prevalence of Bacillus cereus spores during the production of a cooked chilled vegetable product.

In minimally processed vegetable foods, pathogenic spore-forming bacteria pose a significant hazard. As part of a quantitative risk assessment, we used Bayesian belief methods to model the uncertainty and variability of the number of Bacillus cereus spores that can be found in packets of a vegetable puree. The model combines specific information from the manufacturer, experimental data on inactivation of spores, and expert opinion concerning spore concentrations in the raw vegetables and ingredients. Sensitivity analysis revealed that spore contamination of added ingredients contributes most uncertainty to the assessment. The assessment produced a quantitative estimate of the prevalence of B. cereus spores in packets of vegetable puree at the end point of the manufacturing process.

Bacillus cereus↗

Bayesian inference on order-constrained parameters in generalized linear models.

In biomedical studies, there is often interest in assessing the association between one or more ordered categorical predictors and an outcome variable, adjusting for covariates. For a k-level predictor, one typically uses either a k-1 degree of freedom (df) test or a single df trend test, which requires scores for the different levels of the predictor. In the absence of knowledge of a parametric form for the response function, one can incorporate monotonicity constraints to improve the efficiency of tests of association. This article proposes a general Bayesian approach for inference on order-constrained parameters in generalized linear models. Instead of choosing a prior distribution with support on the constrained space, which can result in major computational difficulties, we propose to map draws from an unconstrained posterior density using an isotonic regression transformation. This approach allows flat regions over which increases in the level of a predictor have no effect. Bayes factors for assessing ordered trends can be computed based on the output from a Gibbs sampling algorithm. Results from a simulation study are presented and the approach is applied to data from a time-to-pregnancy study.

Adult↗

A tracking approach to parcellation of the cerebral cortex.

The cerebral cortex is composed of regions with distinct laminar structure. Functional neuroimaging results are often reported with respect to these regions, usually by means of a brain "atlas". Motivated by the need for more precise atlases, and the lack of model-based approaches in prior work in the field, this paper introduces a novel approach to parcellating the cortex into regions of distinct laminar structure, based on the theory of target tracking. The cortical layers are modelled by hidden Markov models and are tracked to determine the Bayesian evidence of layer hypotheses. This model-based parcellation method, evaluated here on a set of histological images of the cortex, is extensible to 3-D images.

Algorithms↗

Variational mixture of Bayesian independent component analyzers.

There has been growing interest in subspace data modeling over the past few years. Methods such as principal component analysis, factor analysis, and independent component analysis have gained in popularity and have found many applications in image modeling, signal processing, and data compression, to name just a few. As applications and computing power grow, more and more sophisticated analyses and meaningful representations are sought. Mixture modeling methods have been proposed for principal and factor analyzers that exploit local gaussian features in the subspace manifolds. Meaningful representations may be lost, however, if these local features are nongaussian or discontinuous. In this article, we propose extending the gaussian analyzers mixture model to an independent component analyzers mixture model. We employ recent developments in variational Bayesian inference and structure determination to construct a novel approach for modeling nongaussian, discontinuous manifolds. We automatically determine the local dimensionality of each manifold and use variational inference to calculate the optimum number of ICA components needed in our mixture model. We demonstrate our framework on complex synthetic data and illustrate its application to real data by decomposing functional magnetic resonance images into meaningful-and medically useful-features.

Bayes Theorem↗

Decoding spike trains instant by instant using order statistics and the mixture-of-Poissons model.

In the brain, spike trains are generated in time and presumably also interpreted as they unfold in time. Recent work (Oram et al., 1999; Baker and Lemon, 2000) suggests that in several areas of the monkey brain, individual spike times carry information because they reflect an underlying rate variation. Constructing a model based on this stochastic structure allows us to apply order statistics to decode spike trains instant by instant as spikes arrive or do not. Order statistics are time-consuming to compute in the general case. We demonstrate that data from neurons in primary visual cortex are well fit by a mixture of Poisson processes; in this special case, our computations are substantially faster. In these data, spike timing contributed information beyond that available from the spike count throughout the trial. At the end of the trial, a decoder based on the mixture-of-Poissons model correctly decoded about three times as many trials as expected by chance, compared with approximately twice as many as expected by chance using the spike count only. If our model perfectly described the spike trains, and enough data were available to estimate model parameters, then our Bayesian decoder would be optimal. For four-fifths of the sets of stimulus-elicited responses, the observed spike trains were consistent with the mixture-of-Poissons model. Most of the error in estimating stimulus probabilities is attributable to not having enough data to specify the parameters of the model rather than to misspecification of the model itself.

Action Potentials↗

Problems in Extinction Model Selection and Parameter Estimation.

It is a vexing problem to achieve a consensus about the proper scientific way to assess population viability for habitat conservation plans. Rather than a hypothesis-testing approach, here it is proposed to select population models, estimate extinction parameters, and assess prediction uncertainty using a pragmatic, empirical Bayesian approach. The simplest usable models include the effects of population growth, r; carrying capacity, K; Allee threshold, N(A); and environmental stochasticity, v(r). Analytic predictions of expected extinction times are available for such models. Models that are more complex can be elaborated from this basis. Selection from a hierarchy of nesting population models can often be done through the evaluation of parameters. The estimation of the most important extinction parameters can be undertaken in a variety of ways. Time series can be analyzed to estimate r(d), v(r), rho, and K. Habitat models and individualistic population models may help estimate N(A) and K and demographic stochasticity. Fine-scale biogeography and climatological data may be useful in the estimation of a variety of parameters. Because it takes many years to estimate extinction parameters accurately for a given population of interest, the most efficient estimation procedures are desirable. I propose the use of prior information from an (as yet nonexistent) population biology database. The accumulation of local information through monitoring will improve our estimates allowing adaptive management. Uncertainty in the estimates will always remain, but it may be quantified by the posterior distributions. A crude example is discussed using treefrog population data. Although the motivations, beliefs, and biases of competing stakeholders will differ, a habitat conservation plan could accommodate this variation in the prior distributions. Field experience from monitoring will increasingly clear up any discrepancies between the opposing beliefs and the real ecosystem. As the world is an uncertain place and because there is no universal scientific method, there will always be controversy and surprises. The best we can do is (1) agree about our prior information, (2) agree about the strategy of model selection and parameter estimation, and (3) agree about our strategy for adaptive management. Perhaps the greatest impediment to such prior agreements for HCPs is the likely paranoia inspired by the use of unfamiliar statistical methodology. We need to train students of ecology in a more flexible and deeper understanding of statistics and philosophy of science.

Journal Article↗

Evaluation of decay times in coupled spaces: Bayesian parameter estimation.

Determination of sound decay times in coupled spaces often demands considerable effort. Based on Schroeder's backward integration of room impulse responses, it is often difficult to distinguish different portions of multirate sound energy decay functions. A model-based parameter estimation method, using Bayesian probabilistic inference, proves to be a powerful tool for evaluating decay times. A decay model due to one of the authors [N. Xiang, J. Acoust. Soc. Am. 98, 2112-2121 (1995)] is extended to multirate decay functions. Following a summary of Bayesian model-based parameter estimation, the present paper discusses estimates in terms of both synthesized and measured decay functions. No careful estimation of initial values is required, in contrast to gradient-based approaches. The resulting robust algorithmic estimation of more than one decay time, from experimentally measured decay functions, is clearly superior to the existing nonlinear regression approach.

Journal Article↗

Theophylline population pharmacokinetics from routine monitoring data in very premature infants with apnoea.

1. Theophylline is commonly used in neonatology for the treatment and prophylaxis of apnoea of prematurity, and during ventilator weaning. 2. NONMEM was used to study the population pharmacokinetics of intravenous and oral theophylline from retrospective drug monitoring data in 82 premature neonates, weighing < 1500 g at birth, and < or = 32 weeks gestational age. 3. Clearance (CL), volume of distribution (V), and oral bioavailability (F1) from liquid preparations were modelled alone, and under the influence of demographic and clinical covariates, assuming a 1-compartment model with first-order elimination. 4. The final population models with influential co-variates were as follows: CL (1h-1) = 0.0000123 *body weight (g) + 0.000377 *postnatal age (days); V (1) = 0.000937 *body weight (g); F = 0.918. 5. The CL was lower and V was higher than previously reported for less premature neonates, term babies, and older children. 6. Predictive performance of the population models was evaluated by Bayesian forecasting in a similar, but independent cohort of 30 infants. There was statistically insignificant bias and imprecision between measured and predicted serum theophylline concentrations. 7. Based on the validated population models, recommended maintenance theophylline dosages are provided for infants aged between 2 and 50 days, and weighing 700 to 2000 g.

Apnea↗

Spatio-temporal autoregressive models defined over brain manifolds.

Multivariate Autoregressive time series models (MAR) are an increasingly used tool for exploring functional connectivity in Neuroimaging. They provide the framework for analyzing the Granger Causality of a given brain region on others. In this article, we shall limit our attention to linear MAR models, in which a set of matrices of autoregressive coefficients Ak (k = 1,...,p) describe the dependence of present values of the image on lagged values of its past. Methods for estimating the Ak and determining which elements that are zero are well-known and are the basis for directed measures of influence. However, to date, MAR models are limited in the number of time series they can handle, forcing the a priori selection of a (small) number of voxels or regions of interest for analysis. This ignores the full spatio-temporal nature of functional brain data which are, in fact, collections of time series sampled over an underlying continuous spatial manifold the brain. A fully spatio-temporal MAR model (ST-MAR) is developed within the framework of functional data analysis. For spatial data, each row of a matrix Ak is the influence field of a given voxel. A Bayesian ST-MAR model is specified in which the influence fields for all voxels are required to vary smoothly over space. This requirement is enforced by penalizing the spatial roughness of the influence fields. This roughness is calculated with a discrete version of the spatial Laplacian operator. A massive reduction in dimensionality of computations is achieved via the singular value decomposition, making an interactive exploration of the model feasible. Use of the model is illustrated with an fMRI time series that was gathered concurrently with EEG in order to analyze the origin of resting brain rhythms.

Bayes Theorem↗

Bayesian classifiers for detecting HGT using fixed and variable order markov models of genomic signatures.

MOTIVATION: Analyses of genomic signatures are gaining attention as they allow studies of species-specific relationships without involving alignments of homologous sequences. A naïve Bayesian classifier was built to discriminate between different bacterial compositions of short oligomers, also known as DNA words. The classifier has proven successful in identifying foreign genes in Neisseria meningitis. In this study we extend the classifier approach using either a fixed higher order Markov model (Mk) or a variable length Markov model (VLMk). RESULTS: We propose a simple algorithm to lock a variable length Markov model to a certain number of parameters and show that the use of Markov models greatly increases the flexibility and accuracy in prediction to that of a naïve model. We also test the integrity of classifiers in terms of false-negatives and give estimates of the minimal sizes of training data. We end the report by proposing a method to reject a false hypothesis of horizontal gene transfer. AVAILABILITY: Software and Supplementary information available at www.cs.chalmers.se/~dalevi/genetic_sign_classifiers/.

Artificial Intelligence↗

Bayesian analysis of response to selection: a case study using litter size in Danish Yorkshire pigs.

Implementation of a Bayesian analysis of a selection experiment is illustrated using litter size [total number of piglets born (TNB)] in Danish Yorkshire pigs. Other traits studied include average litter weight at birth (WTAB) and proportion of piglets born dead (PRBD). Response to selection for TNB was analyzed with a number of models, which differed in their level of hierarchy, in their prior distributions, and in the parametric form of the likelihoods. A model assessment study favored a particular form of an additive genetic model. With this model, the Monte Carlo estimate of the 95% probability interval of response to selection was (0.23; 0.60), with a posterior mean of 0.43 piglets. WTAB showed a correlated response of -7.2 g, with a 95% probability interval equal to (-33.1; 18.9). The posterior mean of the genetic correlation between TNB and WTAB was -0.23 with a 95% probability interval equal to (-0.46; -0.01). PRBD was studied informally; it increases with larger litters, when litter size is >7 piglets born. A number of methodological issues related to the Bayesian model assessment study are discussed, as well as the genetic consequences of inferring response to selection using additive genetic models.

Animal Husbandry↗

A statistical model of transmission of Hib bacteria in a family.

The simultaneous estimation of family and community transmission rates as well as cure rates from panel data in a recurrent Hib (Haemophilus influenzae type b bacteria) infection is considered. An individual-based stationary Markov process model with constant hazards in two age groups is applied to describe recurrent asymptomatic Hib infection in a family with small children. The problem of estimation is solved in terms of the Bayesian posterior of the model parameters. The model is used to predict prevalence and incidence of Hib carriage in families as a function of the family size and age structure.

Adult↗

Maternal animal model with correlation between maternal environmental effects of related dams.

A procedure to take into account the nongenetic relationship between maternal effects in adjacent generations is presented. It considers a correlation between maternal environments provided by a dam and its daughters (lambda). The dispersion structure of the maternal animal model was modified to include a correlation matrix (E) that relates the maternal permanent environmental effects. The structures of the E matrix and its inverse (E(-1)) are described. Both matrices are completely defined by the correlation coefficient lambda. An algorithm to compute these matrices from pedigree information was also developed. Furthermore, a Bayesian analysis of this model including the lambda parameter was developed using Gibbs sampling, with Metropolis steps for the nonstandard conditional distributions. With simulated data, the proposed model reduced the bias in all estimates of dispersion parameters when an antagonism between the maternal effects received by a daughter and its future maternal environment existed. This model also provides an estimate of the environmental relationship between the maternal effects of dams and daughters by the lambda parameter. The same Bayesian analysis was also carried out with weaning weight data of the Bruna dels Pirineus breed. The posterior means (standard deviation) of (co)variance ratios were .214 (.081) for direct heritability (h2d), .107 (.033) for maternal heritability (h2m), .047 (.020) for the proportion of variance due to maternal environmental effects (c2m), and -.034 (.043) for the genetic correlation between direct and maternal effects (r(dm)). The posterior mean of lambda parameter was -.190, and 76% of its marginal posterior distribution took negative values. As occurred with simulated data, considering the maternal environmental correlation in the analysis implied higher h2m estimates, lower c2m and h2d estimates, and less negative values for the marginal posterior distribution of r(dm). These results were considered as evidence of the environmental antagonism between maternal effects provided by a dam and its daughters to weaning weight of their progeny in the Bruna dels Pirineus breed.

Animal Husbandry↗

Bayesian forecasting and prediction of tacrolimus concentrations in pediatric liver and adult renal transplant recipients.

AIM: To test the predictive capacity of two recently derived population pharmacokinetic models and the usefulness of Bayesian forecasting to predict tacrolimus blood concentrations in pediatric liver and adult kidney transplant recipients. MATERIALS AND METHODS: New databases were added to the Abbottbase PKS (Bayesian dosage prediction) program to incorporate the population pharmacokinetic models developed for tacrolimus. Two independent populations of transplant recipients were used to predict tacrolimus trough concentrations. Pharmacokinetic, demographic, and covariate data were collected from patient records. Different time weighting factors were tested (1, 1.005, 1.01) and the influence of excluding data collected in the first 5 days post-transplant examined. Concentrations were predicted until the 10th tacrolimus measurement. Actual tacrolimus concentrations were compared with those predicted by the PKS program and bias and precision determined. RESULTS: Tacrolimus concentrations predicted by the PKS program were, on average, unbiased for the pediatric liver population, but were over-predicted (9%) for the adult renal population. In both populations predictions were not precise (imprecision ranged from 39 to 50%). CONCLUSIONS: Due to the imprecision seen in this study, these models could not be used in clinical practice in the immediate post-transplant period. Poor precision may be due to reliance on routine drug monitoring data alone, difficulties with expression of covariates in continuous modeling relationships in the PKS program, lack of accurate quantitative measures of liver function, or large, random intraindividual variability in the bioavailability of tacrolimus.

Adolescent↗

A comparison of frailty and other models for bivariate survival data.

Multivariate survival data arise when each study subject may experience multiple events or when study subjects are clustered into groups. Statistical analyses of such data need to account for the intra-cluster dependence through appropriate modeling. Frailty models are the most popular for such failure time data. However, there are other approaches which model the dependence structure directly. In this article, we compare the frailty models for bivariate data with the models based on bivariate exponential and Weibull distributions. Bayesian methods provide a convenient paradigm for comparing the two sets of models we consider. Our techniques are illustrated using two examples. One simulated example demonstrates model choice methods developed in this paper and the other example, based on a practical data set of onset of blindness among patients with diabetic Retinopathy, considers Bayesian inference using different models.

Bayes Theorem↗

Contribution of RPB2 to multilocus phylogenetic studies of the euascomycetes (Pezizomycotina, Fungi) with special emphasis on the lichen-forming Acarosporaceae and evolution of polyspory.

Despite the recent progress in molecular phylogenetics, many of the deepest relationships among the main lineages of the largest fungal phylum, Ascomycota, remain unresolved. To increase both resolution and support on a large-scale phylogeny of lichenized and non-lichenized ascomycetes, we combined the protein coding-gene RPB2 with the traditionally used nuclear ribosomal genes SSU and LSU. Our analyses resulted in the naming of the new subclasses Acarosporomycetidae and Ostropomycetidae, and the new class Lichinomycetes, as well as the establishment of the phylogenetic placement and novel circumscription of the lichen-forming fungi family Acarosporaceae. The delimitation of this family has been problematic over the past century, because its main diagnostic feature, true polyspory (numerous spores issued from multiple post-meiosis mitoses) with over 100 spores per ascus, is probably not restricted to the Acarosporaceae. This observation was confirmed by our reconstruction of the origin and evolution of this form of true polyspory using maximum likelihood as the optimality criterion. The various phylogenetic analyses carried out on our data sets allowed us to conclude that: (1) the inclusion of phylogenetic signal from ambiguously aligned regions into the maximum parsimony analyses proved advantageous in reconstructing phylogeny; however, when more data become available, Bayesian analysis using different models of evolution is likely to be more efficient; (2) neighbor-joining bootstrap proportions seem to be more appropriate in detecting topological conflict between data partitions of large-scale phylogenies than posterior probabilities; and (3) Bayesian bootstrap proportion provides a compromise between posterior probability outcomes (i.e., higher accuracy, but with a higher number of significantly supported wrong internodes) vs. maximum likelihood bootstrap proportion outcomes (i.e., lower accuracy, with a lower number of significantly supported wrong internodes).

Ascomycota↗