Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

A Bayesian approach to a general regression model for ROC curves.

A fully Bayesian approach to a general nonlinear ordinal regression model for ROC-curve analysis is presented. Samples from the marginal posterior distributions of the model parameters are obtained by a Markov-chain Monte Carlo (MCMC) technique--Gibbs sampling. These samples facilitate the calculation of point estimates and credible regions as well as inferences for the associated areas under the ROC curves. The analysis of an example using freely available software shows that the use of noninformative vague prior distributions for all model parameters yields posterior summary statistics very similar to the conventional maximum-likelihood estimates. Clinically important advantages of this Bayesian approach are: the possible inclusion of prior knowledge and beliefs into the ROC analysis (via the prior distributions), the possible calculation of the posterior predictive distribution of a future patient outcome, and the potential to address questions such as: "What is the probability that a certain diagnostic test is better in one setting than in another?"

Bayes Theorem↗

Diffusion and prediction of Leishmaniasis in a large metropolitan area in Brazil with a Bayesian space-time model.

We present results from an analysis of human visceral Leishmaniasis cases based on public health records of Belo Horizonte, Brazil, from 1994 to 1997. The main emphasis in this study is on the development of a spatial statistical model to map and project the rates of visceral Leishmaniasis in Belo Horizonte. The model allows for space-time interaction and it is based on a hierarchical Bayesian approach. We assume that the underlying rates evolve in time according to a polynomial trend specific to each small area in the region. The parameters of these polynomials receive a spatial distribution in the form of an autonormal distribution. While the raw rates are extremely noisy and inadequate to support decisions, the resulting smoothed rates estimates are considerably less affected by small area issues and provide very clear directions to implement public health actions.

Bayes Theorem↗

A Bayesian approach to estimate and validate the false negative fraction in a two-stage multiple screening test.

OBJECTIVES: In estimating sensitivity and specificity of a diagnostic kit it is imperative that all study subjects are verified via a gold standard procedure. However the application of such a procedure to all the study subjects may not be feasible due to associated cost, risk and invasiveness. As a result only a part of the study subjects receive the definitive assessment. The accuracy of a diagnostic kit can also be expressed in terms of its error rates. Our first objective is to estimate the false negative fraction (FNF) under partial verification in a particular case of a two-stage multiple screening test using a beta-binomial model and a Bayesian logistic model. The second objective is to validate the two models in order to determine which fits the data better. METHODS: We estimate the FNF from the above mentioned models using Bayesian approach. The validation of the models is based on their out-of-sample predictive capabilities. RESULTS: For the bowel cancer data that was used in this study we found the median posterior estimate of the FNF, based on the beta-binomial model, to be 26.4% (95% credible interval: 0.123-0.650). The corresponding estimate based on the Bayesian logistic model was 23.3% (95% credible interval: 0.124-0.375). Validation results showed that the betabinomial model gave slightly better predictions compared to the Bayesian logistic model. CONCLUSIONS: Estimation of the FNF can be done by adopting the Bayesian approach. Models fitted can be validated by comparing their performance in terms of their out-of-sample predicitve potential.

Bayes Theorem↗

Bayesian analysis of an admixture model with mutations and arbitrarily linked markers.

We introduce here a Bayesian analysis of a classical admixture model in which all parameters are simultaneously estimated. Our approach follows the approximate Bayesian computation (ABC) framework, relying on massive simulations and a rejection-regression algorithm. Although computationally intensive, this approach can easily deal with complex mutation models and partially linked loci, and it can be thoroughly validated without much additional computation cost. Compared to a recent maximum-likelihood (ML) method, the ABC approach leads to similarly accurate estimates of admixture proportions in the case of recent admixture events, but it is found superior when the admixture is more ancient. All other parameters of the admixture model such as the divergence time between parental populations, the admixture time, and the population sizes are also well estimated, unlike the ML method. The use of partially linked markers does not introduce any particular bias in the estimation of admixture, but ML confidence intervals are found too narrow if linkage is not specifically accounted for. The application of our method to an artificially admixed domestic bee population from northwest Italy suggests that the admixture occurred in the last 10-40 generations and that the parental Apis mellifera and A. ligustica populations were completely separated since the last glacial maximum.

Animals↗

A Bayesian population PK-PD model of ispinesib-induced myelosuppression.

The goal of the present analysis is to fit a Bayesian population pharmacokinetic pharmacodynomic (PK-PD) model to characterize the relationship between the concentration of ispinesib and changes in absolute neutrophil counts (ANC). Ispinesib, a kinesin spindle protein (KSP) inhibitor, blocks assembly of a functional mitotic spindle, leading to G2/M arrest. A first time in human, phase I open-label, non-randomized, dose-escalating study evaluated ispinesib at doses ranging from 1 to 21 mg/m(2). PK-PD data were collected from 45 patients with solid tumors. The pharmacokinetics of ispinesib were well characterized by a two-compartment model. A semimechanistic model was fit to the ANC. The PK and PD data were successfully modelled simultaneously. This is the first presentation of simultaneously fitting a PK-PD model to ANC using Bayesian methods. Bayesian methods allow for the use of prior information for some system-related parameters. The model may be used to examine different schedules, doses, and infusion times.

Adult↗

Bayesian and maximum likelihood phylogenetic analyses of protein sequence data under relative branch-length differences and model violation.

BACKGROUND: Bayesian phylogenetic inference holds promise as an alternative to maximum likelihood, particularly for large molecular-sequence data sets. We have investigated the performance of Bayesian inference with empirical and simulated protein-sequence data under conditions of relative branch-length differences and model violation. RESULTS: With empirical protein-sequence data, Bayesian posterior probabilities provide more-generous estimates of subtree reliability than does the nonparametric bootstrap combined with maximum likelihood inference, reaching 100% posterior probability at bootstrap proportions around 80%. With simulated 7-taxon protein-sequence datasets, Bayesian posterior probabilities are somewhat more generous than bootstrap proportions, but do not saturate. Compared with likelihood, Bayesian phylogenetic inference can be as or more robust to relative branch-length differences for datasets of this size, particularly when among-sites rate variation is modeled using a gamma distribution. When the (known) correct model was used to infer trees, Bayesian inference recovered the (known) correct tree in 100% of instances in which one or two branches were up to 20-fold longer than the others. At ratios more extreme than 20-fold, topological accuracy of reconstruction degraded only slowly when only one branch was of relatively greater length, but more rapidly when there were two such branches. Under an incorrect model of sequence change, inaccurate trees were sometimes observed at less extreme branch-length ratios, and (particularly for trees with single long branches) such trees tended to be more inaccurate. The effect of model violation on accuracy of reconstruction for trees with two long branches was more variable, but gamma-corrected Bayesian inference nonetheless yielded more-accurate trees than did either maximum likelihood or uncorrected Bayesian inference across the range of conditions we examined. Assuming an exponential Bayesian prior on branch lengths did not improve, and under certain extreme conditions significantly diminished, performance. The two topology-comparison metrics we employed, edit distance and Robinson-Foulds symmetric distance, yielded different but highly complementary measures of performance. CONCLUSIONS: Our results demonstrate that Bayesian inference can be relatively robust against biologically reasonable levels of relative branch-length differences and model violation, and thus may provide a promising alternative to maximum likelihood for inference of phylogenetic trees from protein-sequence data.

Bayes Theorem↗

Modeling the impact of treatment and screening on U.S. breast cancer mortality: a Bayesian approach.

BACKGROUND: Breast cancer mortality (BCM) in the United States declined from 33.1 per 100,000 women in 1990 to 26.6 per 100,000 women in 2000, yielding a 19.6% relative decline in BCM since 1990. Our goal is to apportion this decline between screening and therapy and to be able to state with some certainty that these interventions affected this decline. METHODS: We started with an age-appropriate population of 2,000,000 women in 1975 and monitored these women through 2000. On the basis of population data each year, we assigned screening and breast cancer to women. If a woman was diagnosed with breast cancer, we simulated a lifetime for her with death from breast cancer, and we modified this lifetime depending on the use of adjuvant therapy and whether the cancer was screen-detected. A woman's lifetime was taken as the minimum of her lifetime with death from breast cancer and her simulated natural lifetime. We used Bayesian simulation modeling, which allows for associating probability distributions with our estimates. RESULTS: We calculated the probabilities that screening mammography and adjuvant therapy contributed to the observed decline in BCM to be 90% and 99%, respectively. The posterior mean reduction in BCM due to screening is 10.6% +/- 5.7% and due to therapy is 19.5% +/- 5.4%. The decrease in the hazard of BCM due to tamoxifen use for ER-positive tumors is 37% +/- 14% and that due to adjuvant (nontaxane) chemotherapy is 15% +/- 14%. DISCUSSION: The spread in our posterior distributions reflect the uncertainty present in the data sources available to us. However, despite this uncertainty we conclude a high probability that both screening and improvements in therapy contributed to the reduction in BCM observed in the United States from 1990 to 2000.

Age Distribution↗

A hierarchical Bayesian approach to age-specific back-calculation of cancer incidence rates.

We propose a Bayesian hierarchical model to estimate age-specific cancer incidence per year from age-specific cancer mortality. The model is based upon the empirical Bayesian approach of Liao and Brookmeyer (1995) and extends that model by consideration of the dependence on age. The incident cases per year are considered as observations from a discrete-time stochastic process following an autoregressive structure within a Poisson regression model. The model assumes that the survival probability among those with cancer is known. We have investigated the sensitivity of the model to the choice of this distribution and have found that this is the most sensitive part of the model. By comparison the predictions of the model are relatively robust to changes in other key areas, such as the number of years an incident case contributes before death, assumptions about parameter equality for identification and the initial prior distributions. The proposed methodology has been investigated using lung cancer mortality data from Scotland. Parameter estimates were obtained through Markov chain Monte Carlo methods, implemented using BUGS.

Age Factors↗

Contrasting Bayesian analysis of survey data and clinical trials.

Although both surveys and clinical trials are amenable to Bayesian hierarchical modelling, the general aims, constraints and actual analysis of each can often vary considerably. First, examples are presented showing how Bayesian hierarchical modelling can be used to produce estimates for small areas from survey data and, also, how it can be used to combine data from clinical trials. Then, it will be shown how surveys and clinical trials may differ with respect to the presence of design effects/selection biases and with the ability to validate models. The impact of the design on modelling will be highlighted and a class of sample selection models will be shown to help alleviate the design's influence. Although surveys generally have enough data to validate many features of a model, clinical trials may not, leaving sensitivity analysis as a means to prior acceptance. Some design issues, contrasting Bayesian with frequentist methods, will also be discussed. Published in 2001 by John Wiley & Sons, Ltd.

Bayes Theorem↗

Detection of mastitis in dairy cattle by use of mixture models for repeated somatic cell scores: a Bayesian approach via Gibbs sampling.

The distribution of somatic cell scores could be regarded as a mixture of at least two components depending on a cow's udder health status. A heteroscedastic two-component Bayesian normal mixture model with random effects was developed and implemented via Gibbs sampling. The model was evaluated using datasets consisting of simulated somatic cell score records. Somatic cell score was simulated as a mixture representing two alternative udder health statuses ("healthy" or "diseased"). Animals were assigned randomly to the two components according to the probability of group membership (Pm). Random effects (additive genetic and permanent environment), when included, had identical distributions across mixture components. Posterior probabilities of putative mastitis were estimated for all observations, and model adequacy was evaluated using measures of sensitivity, specificity, and posterior probability of misclassification. Fitting different residual variances in the two mixture components caused some bias in estimation of parameters. When the components were difficult to disentangle, so were their residual variances, causing bias in estimation of Pm and of location parameters of the two underlying distributions. When all variance components were identical across mixture components, the mixture model analyses returned parameter estimates essentially without bias and with a high degree of precision. Including random effects in the model increased the probability of correct classification substantially. No sizable differences in probability of correct classification were found between models in which a single cow effect (ignoring relationships) was fitted and models where this effect was split into genetic and permanent environmental components, utilizing relationship information. When genetic and permanent environmental effects were fitted, the between-replicate variance of estimates of posterior means was smaller because the model accounted for random genetic drift.

Animals↗

[Lung cancer mortality among women in France. Trend analysis and projection between 1975 and 2014, with a bayesian age-cohort model].

BACKGROUND: With 4,500 deaths in year 2000, female lung cancer mortality rates increased by 3% every year over the last two decades in France. This trend, not observed among males, is attributed to the regular increase of female smoking. In order to answer French Health decider's concerns, we estimated the future female lung cancer mortality rates and numbers of deaths for the next fifteen years, in France and its regions. METHODS: Analyses were based on numbers of female deaths from lung cancer observed between 1975 and 1999, and on past and future population estimates for 1975-2014, at national and regional levels. Mortality rates and numbers of deaths in France and its regions by 5-year periods and 5-year age groups were given in the 1975-1999 death certificate data base, and were projected for 2000-2014. The analysis used a bayesian approach of the age-cohort model with auto-regressive constraints on parameters. Estimated mortality rates were standardized on truncated 20-85 + world population. RESULTS: French female lung cancer mortality increased by 3% every year between 1975 and 1999. In period 1995-1999, truncated 20-85 + mortality rates, and number of deaths per year were respectively 11.4 per 100,000 and 4,000. Mortality rates increased in all regions but variations were maximum in Corsica (+ 314%) and minimum in Auvergne (+ 37%). For the whole of France, the estimated truncated 20-85 + standardized rate, was respectively, 14.1 and 22.5 per 100,000 in period 2000-04 and period 2010-14, which represents a 60% increase between these two time periods. At the regional level, the maximum variation was found in Languedoc-Roussillon (107%), the minimum in Nord-Pas-de-Calais (40%). CONCLUSIONS: The bayesian approach of the age-cohort model is increasingly used because it produces stable projections, without having to include other cancer parameters. Nevertheless, it would be interesting to extend this model by incorporating a tobacco consumption component, in order to assess scenarios based on consumption decreases.

Adult↗

A model-based approach to Bayesian classification with applications to predicting pregnancy outcomes from longitudinal beta-hCG profiles.

This paper discusses Bayesian statistical methods for the classification of observations into two or more groups based on hierarchical models for nonlinear longitudinal profiles. Parameter estimation for a discriminant model that classifies individuals into distinct predefined groups or populations uses appropriate posterior simulation schemes. The methods are illustrated with data from a study involving 173 pregnant women. The main objective in this study is to predict normal versus abnormal pregnancy outcomes from beta human chorionic gonadotropin data available at early stages of pregnancy.

Bayes Theorem↗

Using literature and data to learn Bayesian networks as clinical models of ovarian tumors.

Thanks to its increasing availability, electronic literature has become a potential source of information for the development of complex Bayesian networks (BN), when human expertise is missing or data is scarce or contains much noise. This opportunity raises the question of how to integrate information from free-text resources with statistical data in learning Bayesian networks. Firstly, we report on the collection of prior information resources in the ovarian cancer domain, which includes "kernel" annotations of the domain variables. We introduce methods based on the annotations and literature to derive informative pairwise dependency measures, which are derived from the statistical cooccurrence of the names of the variables, from the similarity of the "kernel" descriptions of the variables and from a combined method. We perform wide-scale evaluation of these text-based dependency scores against an expert reference and against data scores (the mutual information (MI) and a Bayesian score). Next, we transform the text-based dependency measures into informative text-based priors for Bayesian network structures. Finally, we report the benefit of such informative text-based priors on the performance of a Bayesian network for the classification of ovarian tumors from clinical data.

Artificial Intelligence↗

Bayesian analysis of binary prediction tree models for retrospectively sampled outcomes.

Classification tree models are flexible analysis tools which have the ability to evaluate interactions among predictors as well as generate predictions for responses of interest. We describe Bayesian analysis of a specific class of tree models in which binary response data arise from a retrospective case-control design. We are also particularly interested in problems with potentially very many candidate predictors. This scenario is common in studies concerning gene expression data, which is a key motivating example context. Innovations here include the introduction of tree models that explicitly address and incorporate the retrospective design, and the use of nonparametric Bayesian models involving Dirichlet process priors on the distributions of predictor variables. The model specification influences the generation of trees through Bayes' factor based tests of association that determine significant binary partitions of nodes during a process of forward generation of trees. We describe this constructive process and discuss questions of generating and combining multiple trees via Bayesian model averaging for prediction. Additional discussion of parameter selection and sensitivity is given in the context of an example which concerns prediction of breast tumour status utilizing high-dimensional gene expression data; the example demonstrates the exploratory/explanatory uses of such models as well as their primary utility in prediction. Shortcomings of the approach and comparison with alternative tree modelling algorithms are also discussed, as are issues of modelling and computational extensions.

Algorithms↗

Bayesian meta-analysis, with application to studies of ETS and lung cancer.

Meta-analysis enables researchers to combine the results of several studies to assess the information they provide as a whole. It has been used to give a systematic overview of many areas in which data on a possible association between an exposure and an outcome have been collected in a number of studies but where the overall picture remains obscure, both as to the existence or size of the effect. This paper outlines some innovations in meta-analysis, based on using Markov chain Monte Carlo (MCMC) techniques for implementing Bayesian hierarchical models, and compares these with a more well-known random effects (RE) model. The new techniques allow different aspects of variation to be incorporated into descriptions of the association, and in particular enable researchers to better quantify differences between studies. Both the classical and Bayesian methods are applied, in this paper, to the current collection of studies of the association between incidence of lung cancer in female never-smokers and exposure to environmental tobacco smoke (ETS), both in the home through spousal smoking and in the workplace. In this paper it is demonstrated that compared with the RE model, the Bayesian methods: (a) allow more detailed modeling of study heterogeneity to be incorporated; (b) are relatively robust against a wide choice of specifications of such information on heterogeneity; (c) allow for more detailed and satisfactory statements to be made, not only about the overall risk but about the individual studies, on the basis of the combined information. For the workplace exposure data set, the Bayesian methods give a somewhat lower overall estimate of relative risk of lung cancer associated with ETS, indicating the care that needs to be taken in using point estimates based on any one method of analysis. On the larger spousal data set the methods give similar answers. Some of the other concerns with meta-analysis are also considered. These include: consistency between different geographic areas (Asia and the United States), and our studies show that Bayesian methods permit an account of the overall picture to be taken, thus improving the ability to estimate accurately in the subgroups; and publication bias which, as shown with the spousal exposure data, may lead to an inflated excess risk.

Air Pollution, Indoor↗

Uncertainty Modeling Outperforms Machine Learning for Microbiome Data Analysis.

Microbiome sequencing measures relative rather than absolute abundances, providing no direct information about total microbial load. Normalization methods attempt to compensate, but rely on strong, often untestable assumptions that can bias inference. Experimental measurements of load (e.g., qPCR, flow cytometry) offer a solution, but remain costly and uncommon. A recent high-profile study proposed that machine learning could bypass this limitation by predicting microbial load from sequencing data alone. To evaluate this claim, we assembled mutt, the largest public database of paired sequencing and load measurements, spanning 35 studies and over 15,000 samples. Using mutt, we show that published machine learning models fail to generalize: on average they perform worse than a naive baseline that always predicted the training set mean. These failures stem from covariate shift-limited shared taxa between studies, differences in community composition, and differences in preprocessing pipelines-that silently derail model inputs. In contrast, Bayesian partially identified models do not attempt to impute microbial load, but instead propagate scale uncertainty through downstream analyses. Across 30 benchmark datasets, Bayesian partially identified models consistently outperformed normalization and machine learning approaches, providing a principled and reproducible foundation for microbiome inference.

16S rRNA-seq↗

Impact of nonignorable coarsening on Bayesian inference.

The coarse data model of Heitjan and Rubin (1991) generalizes the missing data model of Rubin (1976) to cover other forms of incompleteness such as censoring and grouping. The model has 2 components: an ideal data model describing the distribution of the quantity of interest and a coarsening mechanism that describes a distribution over degrees of coarsening given the ideal data. The coarsening mechanism is said to be nonignorable when the degree of coarsening depends on an incompletely observed ideal outcome, in which case failure to properly account for it can spoil inferences. A theme in recent research is to measure sensitivity to nonignorability by evaluating the effect of a small departure from ignorability on the maximum likelihood estimate (MLE) of a parameter of the ideal data model. One such construct is the "index of local sensitivity to nonignorability" (ISNI) (Troxel and others, 2004), which is the derivative of the MLE with respect to a nonignorability parameter evaluated at the ignorable model. In this paper, we adapt ISNI to Bayesian modeling by instead defining it as the derivative of the posterior expectation. We propose the application of ISNI as a first step in judging the robustness of a Bayesian analysis to nonignorable coarsening. We derive formulas for a range of models and apply the method to evaluate sensitivity to nonignorable coarsening in 2 real data examples, one involving missing CD4 counts in an HIV trial and the other involving potentially informatively censored relapse times in a leukemia trial.

Bayes Theorem↗