Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Phylogenetic systematics turns over a new leaf.

Long restricted to the domain of molecular systematics and studies of molecular evolution, likelihood methods are now being used in analyses of discrete morphological data, specifically to estimate ancestral character states and for tests of character correlation. Biologists are beginning to apply likelihood models within a Bayesian statistical framework, which promises not only to provide answers that evolutionary biologists desire, but also to make practical the application of more realistic evolutionary models.

Journal Article↗

Reliability of Bayesian posterior probabilities and bootstrap frequencies in phylogenetics.

Many empirical studies have revealed considerable differences between nonparametric bootstrapping and Bayesian posterior probabilities in terms of the support values for branches, despite claimed predictions about their approximate equivalence. We investigated this problem by simulating data, which were then analyzed by maximum likelihood bootstrapping and Bayesian phylogenetic analysis using identical models and reoptimization of parameter values. We show that Bayesian posterior probabilities are significantly higher than corresponding nonparametric bootstrap frequencies for true clades, but also that erroneous conclusions will be made more often. These errors are strongly accentuated when the models used for analyses are underparameterized. When data are analyzed under the correct model, nonparametric bootstrapping is conservative. Bayesian posterior probabilities are also conservative in this respect, but less so.

Bayes Theorem↗

Bayesian inference applied to the electromagnetic inverse problem.

We present a new approach to the electromagnetic inverse problem that explicitly addresses the ambiguity associated with its ill-posed character. Rather than calculating a single "best" solution according to some criterion, our approach produces a large number of likely solutions that both fit the data and any prior information that is used. Whereas the range of the different likely results is representative of the ambiguity in the inverse problem even with prior information present, features that are common across a large number of the different solutions can be identified and are associated with a high degree of probability. This approach is implemented and quantified within the formalism of Bayesian inference, which combines prior information with that of measurement in a common framework using a single measure. To demonstrate this approach, a general neural activation model is constructed that includes a variable number of extended regions of activation and can incorporate a great deal of prior information on neural current such as information on location, orientation, strength, and spatial smoothness. Taken together, this activation model and the Bayesian inferential approach yield estimates of the probability distributions for the number, location, and extent of active regions. Both simulated MEG data and data from a visual evoked response experiment are used to demonstrate the capabilities of this approach.

Bayes Theorem↗

Interval estimation of urban ozone level and selection of influential factors by employing automatic relevance determination model.

In this work, we focus on simulating the ground-level ozone (O3) time series and its daily maximum concentration in Hong Kong urban air by employing the multilayer perceptron (MLP) model combined with the automatic relevance determination (ARD) method (for simplicity, we name it as MLP-ARD model). Two air quality monitoring sites in Hong Kong, i.e., Tsuen Wan and Tung Chung, are selected for the numerical experiments. The MLP-ARD model based on Bayesian evidence framework can provide reliable interval estimation of real observation as well as offering efficient strategy to avoid over-fitting. The performance comparisons between MLP-ARD model and traditional artificial neural network (ANN) model based on maximum likelihood indicate that MLP-ARD model is more powerful to capture the wild fluctuation of O3 level especially during O3 episodes than the traditional model. Furthermore, it can assess and rank the input variables for the prediction according to their relative importance to the output variable, i.e., the daily maximum O3 concentration in this study. The preliminary experimental results indicate that nitric oxide (NO) and solar radiation are the most important input variables for O3 prediction at both selected sites. In addition, the previous daily maximum O3 level is also important for Tung Chung site. In this regard, MLP-ARD model is a feasible tool to interpret the real physical and chemical process of urban O3 variation.

Air Pollutants↗

Accuracy of rate estimation using relaxed-clock models with a critical focus on the early metazoan radiation.

In recent years, a number of phylogenetic methods have been developed for estimating molecular rates and divergence dates under models that relax the molecular clock constraint by allowing rate change throughout the tree. These methods are being used with increasing frequency, but there have been few studies into their accuracy. We tested the accuracy of several relaxed-clock methods (penalized likelihood and Bayesian inference using various models of rate change) using nucleotide sequences simulated on a nine-taxon tree. When the sequences evolved with a constant rate, the methods were able to infer rates accurately, but estimates were more precise when a molecular clock was assumed. When the sequences evolved under a model of auto-correlated rate change, rates were accurately estimated using penalized likelihood and by Bayesian inference using lognormal and exponential models of rate change, while other models did not perform as well. When the sequences evolved under a model of uncorrelated rate change, only Bayesian inference using an exponential rate model performed well. Collectively, the results provide a strong recommendation for using the exponential model of rate change if a conservative approach to divergence time estimation is required. A case study is presented in which we use a simulation-based approach to examine the hypothesis of elevated rates in the Cambrian period, and it is found that these high rate estimates might be an artifact of the rate estimation method. If this bias is present, then the ages of metazoan divergences would be systematically underestimated. The results of this study have implications for studies of molecular rates and divergence dates.

Algorithms↗

Bayesian estimation of cerebral perfusion using a physiological model of microvasculature.

Perfusion weighted MRI has proven very useful for deriving hemodynamic parameters such as CBF, CBV and MTT. These quantities are important diagnostically, e.g. in acute stroke, where they are used to delineate ischemic regions. Yet the standard method for estimating CBF based on singular value decomposition (SVD) has been demonstrated to underestimate (especially high) flow components and to be sensitive to delays in the arterial input function (AIF). Furthermore, the estimated residue functions often oscillate. This compromises their physiological interpretation/basis and makes estimation of related measures such as flow heterogeneity difficult. In this study, we estimate perfusion parameters based on a vascular model (VM) which represents heterogeneous capillary flow and explicitly leads to monotonically decreasing residue functions. We use a fully Bayesian approach to obtain posterior probability distributions for all parameters. In simulation studies, we show that the VM method has less bias in CBF estimates than the SVD based method for realistic SNRs. This also applies to cases where the AIF is delayed. We employ our method to estimate perfusion maps using data from (i) a healthy volunteer and (ii) from a stroke patient.

Bayes Theorem↗

Hierarchical phylogenetic models for analyzing multipartite sequence data.

Debate exists over how to incorporate information from multipartite sequence data in phylogenetic analyses. Strict combined-data approaches argue for concatenation of all partitions and estimation of one evolutionary history, maximizing the explanatory power of the data. Consensus/independence approaches endorse a two-step procedure where partitions are analyzed independently and then a consensus is determined from the multiple results. Mixtures across the model space of a strict combined-data approach and a priori independent parameters are popular methods to integrate these methods. We propose an alternative middle ground by constructing a Bayesian hierarchical phylogenetic model. Our hierarchical framework enables researchers to pool information across data partitions to improve estimate precision in individual partitions while permitting estimation and testing of tendencies in across-partition quantities. Such across-partition quantities include the distribution from which individual topologies relating the sequences within a partition are drawn. We propose standard hierarchical priors on continuous evolutionary parameters across partitions, while the structure on topologies varies depending on the research problem. We illustrate our model with three examples. We first explore the evolutionary history of the guinea pig (Cavia porcellus) using alignments of 13 mitochondrial genes. The hierarchical model returns substantially more precise continuous parameter estimates than an independent parameter approach without losing the salient features of the data. Second, we analyze the frequency of horizontal gene transfer using 50 prokaryotic genes. We assume an unknown species-level topology and allow individual gene topologies to differ from this with a small estimable probability. Simultaneously inferring the species and individual gene topologies returns a transfer frequency of 17%. We also examine HIV sequences longitudinally sampled from HIV+ patients. We ask whether posttreatment development of CCR5 coreceptor virus represents concerted evolution from middisease CXCR4 virus or reemergence of initial infecting CCR5 virus. The hierarchical model pools partitions from multiple unrelated patients by assuming that the topology for each patient is drawn from a multinomial distribution with unknown probabilities. Preliminary results suggest evolution and not reemergence.

Animals↗

Patterns of gene expression that characterize long-term survival in advanced stage serous ovarian cancers.

PURPOSE: A better understanding of the underlying biology of invasive serous ovarian cancer is critical for the development of early detection strategies and new therapeutics. The objective of this study was to define gene expression patterns associated with favorable survival. EXPERIMENTAL DESIGN: RNA from 65 serous ovarian cancers was analyzed using Affymetrix U133A microarrays. This included 54 stage III/IV cases (30 short-term survivors who lived <3 years and 24 long-term survivors who lived >7 years) and 11 stage I/II cases. Genes were screened on the basis of their level of and variability in expression, leaving 7,821 for use in developing a predictive model for survival. A composite predictive model was developed that combines Bayesian classification tree and multivariate discriminant models. Leave-one-out cross-validation was used to select and evaluate models. RESULTS: Patterns of genes were identified that distinguish short-term and long-term ovarian cancer survivors. The expression model developed for advanced stage disease classified all 11 early-stage ovarian cancers as long-term survivors. The MAL gene, which has been shown to confer resistance to cancer therapy, was most highly overexpressed in short-term survivors (3-fold compared with long-term survivors, and 29-fold compared with early-stage cases). These results suggest that gene expression patterns underlie differences in outcome, and an examination of the genes that provide this discrimination reveals that many are implicated in processes that define the malignant phenotype. CONCLUSIONS: Differences in survival of advanced ovarian cancers are reflected by distinct patterns of gene expression. This biological distinction is further emphasized by the finding that early-stage cancers share expression patterns with the advanced stage long-term survivors, suggesting a shared favorable biology.

Case-Control Studies↗

Modeling species-habitat relationships with spatially autocorrelated observation data.

Spatial autocorrelation in wildlife observation data arises when extrinsic environmental processes and patterns that influence the spatial distribution of wildlife are themselves spatially structured, or when species are subject to intrinsic population processes, causing contagion or dispersion effects. Territoriality, Allee effects, dispersal limitations, and social clustering are examples of intrinsic processes. Both forms of autocorrelation can violate the assumptions of generalized linear regression models, resulting in biased estimation of model coefficients and diminished predictive performance. Such consequences may be avoided for extrinsic autocorrelation when autocorrelated environmental variables are available for use as model covariates, whereas intrinsic spatial autocorrelation requires an alternative modeling approach. The autologistic model provides an approach suited to the binary observations often obtained in wildlife surveys, but its performance has not been tested across widely varying sampling intensities or strengths of intrinsic spatial structure. Here we use simulated data to test the autologistic model under a range of sampling conditions. The autologistic model obtains better fits and substantially better predictive performance than the standard logistic regression model over the full range of sampling designs and intensities tested. We provide a simple Bayesian implementation of the autologistic model, which until now has not been achieved with standard statistical software alone. A step-by-step procedure is given for characterizing and modeling spatial autocorrelation in binary observation data, along with computer code for fitting autologistic models in WinBUGS, a freeware Bayesian analysis package. This approach avoids normal approximations to the pseudo-likelihood, in contrast to previous Bayesian applications of the autologistic model. We provide a sample application of the autologistic model, fitted to survey data for a gliding marsupial in southeastern Australia.

Bayes Theorem↗

A heuristic Bayesian method for segmenting DNA sequence alignments and detecting evidence for recombination and gene conversion.

We propose a heuristic approach to the detection of evidence for recombination and gene conversion in multiple DNA sequence alignments. The proposed method consists of two stages. In the first stage, a sliding window is moved along the DNA sequence alignment, and phylogenetic trees are sampled from the conditional posterior distribution with MCMC. To reduce the noise intrinsic to inference from the limited amount of data available in the typically short sliding window, a clustering algorithm based on the Robinson-Foulds distance is applied to the trees thus sampled, and the posterior distribution over tree clusters is obtained for each window position. While changes in this posterior distribution are indicative of recombination or gene conversion events, it is difficult to decide when such a change is statistically significant. This problem is addressed in the second stage of the proposed algorithm, where the distributions obtained in the first stage are post-processed with a Bayesian hidden Markov model (HMM). The emission states of the HMM are associated with posterior distributions over phylogenetic tree topology clusters. The hidden states of the HMM indicate putative recombinant segments. Inference is done in a Bayesian sense, sampling parameters from the posterior distribution with MCMC. Of particular interest is the determination of the number of hidden states as an indication of the number of putative recombinant regions. To this end, we apply reversible jump MCMC, and sample the number of hidden states from the respective posterior distribution.

Actins↗

Bayesian analyses of longitudinal binary data using Markov regression models of unknown order.

We present non-homogeneous Markov regression models of unknown order as a means to assess the duration of autoregressive dependence in longitudinal binary data. We describe a subject's transition probability evolving over time using logistic regression models for his or her past outcomes and covariates. When the initial values of the binary process are unknown, they are treated as latent variables. The unknown initial values, model parameters, and the order of transitions are then estimated using a Bayesian variable selection approach, via Gibbs sampling. As a comparison with our approach, we also implement the deviance information criterion (DIC) for the determination of the order of transitions. An example addresses the progression of substance use in a community sample of n = 242 American Indian children who were interviewed annually four times. An extension of the Markov model to account for subject-to-subject heterogeneity is also discussed.

Adolescent↗

Bayesian selection of continuous-time Markov chain evolutionary models.

We develop a reversible jump Markov chain Monte Carlo approach to estimating the posterior distribution of phylogenies based on aligned DNA/RNA sequences under several hierarchical evolutionary models. Using a proper, yet nontruncated and uninformative prior, we demonstrate the advantages of the Bayesian approach to hypothesis testing and estimation in phylogenetics by comparing different models for the infinitesimal rates of change among nucleotides, for the number of rate classes, and for the relationships among branch lengths. We compare the relative probabilities of these models and the appropriateness of a molecular clock using Bayes factors. Our most general model, first proposed by Tamura and Nei, parameterizes the infinitesimal change probabilities among nucleotides (A, G, C, T/U) into six parameters, consisting of three parameters for the nucleotide stationary distribution, two rate parameters for nucleotide transitions, and another parameter for nucleotide transversions. Nested models include the Hasegawa, Kishino, and Yano model with equal transition rates and the Kimura model with a uniform stationary distribution and equal transition rates. To illustrate our methods, we examine simulated data, 16S rRNA sequences from 15 contemporary eubacteria, halobacteria, eocytes, and eukaryotes, 9 primates, and the entire HIV genome of 11 isolates. We find that the Kimura model is too restrictive, that the Hasegawa, Kishino, and Yano model can be rejected for some data sets, that there is evidence for more than one rate class and a molecular clock among similar taxa, and that a molecular clock can be rejected for more distantly related taxa.

Animals↗

HMO selection and Medicare costs: Bayesian MCMC estimation of a robust panel data tobit model with survival.

The fraction of US Medicare recipients enrolled in health maintenance organizations (HMOs) has increased substantially over the past 10 years. However, the impact of HMOs on health care costs is still hotly debated. In particular, it is argued that HMOs achieve cost reduction through 'cream-skimming' and enrolling relatively healthy patients. This paper develops a Bayesian panel data tobit model of HMO selection and Medicare expenditures for recent US retirees that accounts for mortality over the course of the panel. The model is estimated using Markov Chain Monte Carlo (MCMC) simulation methods, and is novel in that a multivariate t-link is used in place of normality to allow for the heavy-tailed distributions often found in health care expenditure data. The findings indicate that HMOs select individuals who are less likely to have positive health care expenditures prior to enrollment. However, there is no evidence that HMOs disenrol high cost patients. The results also indicate the importance of accounting for survival over the panel, since high mortality probabilities are associated with higher health care expenditures in the last year of life.

Aged↗

Functional form and risk adjustment of hospital costs: Bayesian analysis of a Box-Cox random coefficients model.

While risk-adjusted outcomes are often used to compare the performance of hospitals and physicians, the most appropriate functional form for the risk adjustment process is not always obvious for continuous outcomes such as costs. Semi-log models are used most often to correct skewness in cost data, but there has been limited research to determine whether the log transformation is sufficient or whether another transformation is more appropriate. This study explores the most appropriate functional form for risk-adjusting the cost of coronary artery bypass graft (CABG) surgery. Data included patients undergoing CABG surgery at four hospitals in the midwest and were fit to a Box-Cox model with random coefficients (BCRC) using Markov chain Monte Carlo methods. Marginal likelihoods and Bayes factors were computed to perform model comparison of alternative model specifications. Rankings of hospital performance were created from the simulation output and the rankings produced by Bayesian estimates were compared to rankings produced by standard models fit using classical methods. Results suggest that, for these data, the most appropriate functional form is not logarithmic, but corresponds to a Box-Cox transformation of -1. Furthermore, Bayes factors overwhelmingly rejected the natural log transformation. However, the hospital ranking induced by the BCRC model was not different from the ranking produced by maximum likelihood estimates of either the linear or semi-log model.

Bayes Theorem↗

Pharmaco-informatics: more precise drug therapy from "multiple model" (MM) stochastic adaptive control regimens: evaluation with simulated vancomycin therapy.

MM stochastic control of dosage regimens permits essentially full use of information, either in a population pharmacokinetic model or a Bayesian updated MM parameter set, to achieve and maintain selected therapeutic goals with optimal precision. The regimens are visibly more precise than those developed using mean parameter values. Bayesian MM feedback has now also been implemented.

Bayes Theorem↗

Model-free analysis of protein dynamics: assessment of accuracy and model selection protocols based on molecular dynamics simulation.

The popular model-free approach to analyze NMR relaxation measurements has been examined using artificial amide (15)N relaxation data sets generated from a 10 nanosecond molecular dynamics trajectory of a dihydrofolate reductase ternary complex in explicit water. With access to a detailed picture of the underlying internal motions, the efficacy of model-free analysis and impact of model selection protocols on the interpretation of NMR data can be studied. In the limit of uncorrelated global tumbling and internal motions, fitting the relaxation data to the model-free models can recover a significant amount of quantitative information on the internal dynamics. Despite a slight overestimation, the generalized order parameter is quite accurately determined. However, the model-free analysis appears to be insensitive to the presence of nanosecond time scale motions with relatively small magnitude. For such cases, the effective correlation time can be significantly underestimated. As a result, proteins appear to be more rigid than they really are. The model selection protocols have a major impact on the information one can reliably obtain. The commonly employed protocol based on step-up hypothesis testing has severe drawbacks of oversimplification and underfitting. The consequences are that the order parameter is more severely overestimated and the correlation time more severely underestimated. Instead, model selection based on Bayesian Information Criteria (BIC), recently introduced to the model-free analysis by d'Auvergne and Gooley (2003), provides a better balance between bias and variance. More appropriate models can be selected, leading to improved estimate of both the order parameter and correlation time. In addition, the computational cost is significantly reduced and subjective parameters such as the significance level are unnecessary.

Anisotropy↗

Errors-in-variables in joint population pharmacokinetic/pharmacodynamic modeling.

Pharmacokinetic (PK) models describe the relationship between the administered dose and the concentration of drug (and/or metabolite) in the blood as a function of time. Pharmacodynamic (PD) models describe the relationship between the concentration in the blood (or the dose) and the biologic response. Population PK/PD studies aim to determine the sources of variability in the observed concentrations/responses across groups of individuals. In this article, we consider the joint modeling of PK/PD data. The natural approach is to specify a joint model in which the concentration and response data are simultaneously modeled. Unfortunately, this approach may not be optimal if, due to sparsity of concentration data, an overly simple PK model is specified. As an alternative, we propose an errors-in-variables approach in which the observed-concentration data are assumed to be measured with error without reference to a specific PK model. We give an example of an analysis of PK/PD data obtained following administration of an anticoagulant drug. The study was originally carried out in order to make dosage recommendations. The prior for the distribution of the true concentrations, which may incorporate an individual's covariate information, is derived as a predictive distribution from an earlier study. The errors-in-variables approach is compared with the joint modeling approach and more naive methods in which the observed concentrations, or the separately modeled concentrations, are substituted into the response model. Throughout, a Bayesian approach is taken with implementation via Markov chain Monte Carlo methods.

Anticoagulants↗

Using Dirichlet mixture priors to derive hidden Markov models for protein families.

A Bayesian method for estimating the amino acid distributions in the states of a hidden Markov model (HMM) for a protein family or the columns of a multiple alignment of that family is introduced. This method uses Dirichlet mixture densities as priors over amino acid distributions. These mixture densities are determined from examination of previously constructed HMMs or multiple alignments. It is shown that this Bayesian method can improve the quality of HMMs produced from small training sets. Specific experiments on the EF-hand motif are reported, for which these priors are shown to produce HMMs with higher likelihood on unseen data, and fewer false positives and false negatives in a database search task.

Amino Acid Sequence↗