Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,729 records · Page 96Linked to original sources

A chain of evidence with mixed comparisons: models for multi-parameter synthesis and consistency of evidence.

Multi-parameter evidence synthesis is a generalization of meta-analysis in which several parameters are estimated jointly. Here this approach is applied to a common form of data structure in which mixed pairwise comparisons are made between treatments in a 'chain of evidence' structure. The data set investigated here, which first appeared in the confidence profile method literature, features three types of data relating to thrombolytic treatment following acute myocardial infarction. There is information on reperfusion (coronary patency) following treatment, information on survival in reperfused and non-reperfused patients (conditional survival), and independent information on the effect of treatment on survival without specifying reperfusion state (overall survival). The objective of this study is to explore models for combining these three types of evidence within a single model, and some of the evidence consistency issues that arise. Bayesian Markov chain Monte Carlo methods are used to fit two models. The first, proposed in the confidence profile literature, assumes there are no differences between studies in baseline effects (equal study effects). The second model assumes fixed treatment effects but allows for study differences in baseline reperfusion and conditional survival rates by assuming random study effects. The equal study effects model fits the data poorly, and cross-validation shows that the overall survival evidence is not consistent with the reperfusion and conditional survival evidence under this model. The evidence appears to be consistent within the more loosely structured random study effects model, but results are sensitive to prior assumptions about the between-study variance. The inconsistency between the three evidence types relates to overall survival, rather than to the effect of treatment on survival. Multi-parameter evidence synthesis, with suitable checks for evidence consistency, can be used to help determine whether chains of evidence 'add up' epidemiologically within an overall model, and to construct parameter distributions for an entire model simultaneously, consistent with all the available data.

Bayes Theorem↗

Statistical approach to neural network model building for gentamicin peak predictions.

Feed forward neural networks are flexible, nonlinear modeling tools that are an extension of traditional statistical techniques. The hypothesis that feed forward neural network models can be built in a similar fashion as a statistical model was tested. Feed forward neural network models were built using forward and backward variable selection, and zero to five hidden nodes, and tanh and linear transfer functions were used. Gentamicin serum concentrations were predicted as a model drug for testing these methods. Peak observations from 392 patients were used to train, test, and validate the feed forward neural network. Inputs were demographic and drug dosing information. Model selection was performed using the Akaike information criteria (AIC), Bayesian information criteria (BIC), and a method of stopped training. The models with lowest root mean square (rms) error were those with all 10 inputs and five hidden nodes. Average rms error in the validation set was lowest for stopped training (1.46), then AIC (1.51), and finally BIC (1.56). Larger models tended to result in the best predictions. Overfitting can occur in models that are too large, either by using too many nodes in the hidden layer (rms = 1.49) or by using too many inputs with little information associated with them (rms = 1.70). We conclude that neural networks can be built using a large number of parameters that have good predictive performance. Care must be used during training to avoid overfitting the data. A stopped training method resulted in the network with the lowest rms error.

Adult↗

A mixture model for longitudinal data with application to assessment of noncompliance.

In clinical trials of a self-administered drug, repeated measures of a laboratory marker, which is affected by study medication and collected in all treatment arms, can provide valuable information on population and individual summaries of compliance. In this paper, we introduce a general finite mixture of nonlinear hierarchical models that allows estimates of component membership probabilities and random effect distributions for longitudinal data arising from multiple subpopulations, such as from noncomplying and complying subgroups in clinical trials. We outline a sampling strategy for fitting these models, which consists of a sequence of Gibbs, Metropolis-Hastings, and reversible jump steps, where the latter is required for switching between component models of different dimensions. Our model is applied to identify noncomplying subjects in the placebo arm of a clinical trial assessing the effectiveness of zidovudine (AZT) in the treatment of patients with HIV, where noncompliance was defined as initiation of AZT during the trial without the investigators' knowledge. We fit a hierarchical nonlinear change-point model for increases in the marker MCV (mean corpuscular volume of erythrocytes) for subjects who noncomply and a constant mean random effects model for those who comply. As part of our fully Bayesian analysis, we assess the sensitivity of conclusions to prior and modeling assumptions and demonstrate how external information and covariates can be incorporated to distinguish subgroups.

Anti-HIV Agents↗

A four-shock Bayesian up-down estimator of the 80% effective defibrillation dose.

INTRODUCTION: New defibrillation techniques are often compared to standard approaches using the defibrillation threshold. However, inference from thresholding data necessitates extrapolation from reactions to relatively ineffective shocks, an error prone procedure requiring large sample sizes for hypothesis testing and large safety margins for defibrillator implantation. In contrast, this article presents a clinically validated statistical model of a minimum error, four-shock defibrillation testing protocol for estimating the 80% effective defibrillation strength for a given patient (ED80). METHODS AND RESULTS: A Bayesian statistical model was constructed assuming that the defibrillation dose-response curve is sigmoidal, and the ED80 is between 150 and 750 V. The model was used to design a minimum predicted error testing protocol and estimates. To prospectively validate the testing protocol and estimates, 170 patients received voltage-programmed biphasic testing. Four fibrillation episodes were induced and terminated in each patient according to the Bayesian up-down protocol. In addition, a validation attempt was made at the estimated ED80 rounded up to the nearest 50 V. In order to estimate the safety margin, in 136 patients, a defibrillation attempt was made at the rounded ED80 + 100 V. Of the 170 attempts at the rounded ED80, 143 (84%) attempts terminated fibrillation. Of the 136 attempts at the rounded ED80 + 100 V, 133 (98%) were effective. CONCLUSIONS: The four-shock Bayesian up-down protocol is the first clinical protocol to accurately predict an ED80 voltage. A 100 V increment above the ED80 provides an adequate safety margin. This simple and accurate method for estimating a highly effective defibrillation dose may be a valuable tool for population-based clinical hypothesis testing, as well as defibrillator implantation.

Adult↗

A Bayesian hierarchical approach for relating PM(2.5) exposure to cardiovascular mortality in North Carolina.

Considerable attention has been given to the relationship between levels of fine particulate matter (particulate matter < or = 2.5 microm in aerodynamic diameter; PM(2.5) in the atmosphere and health effects in human populations. Since the U.S. Environmental Protection Agency began widespread monitoring of PM(2.5) levels in 1999, the epidemiologic community has performed numerous observational studies modeling mortality and morbidity responses to PM(2.5) levels using Poisson generalized additive models (GAMs). Although these models are useful for relating ambient PM(2.5) levels to mortality, they cannot directly measure the strength of the effect of exposure to PM(2.5) on mortality. In order to assess this effect, we propose a three-stage Bayesian hierarchical model as an alternative to the classical Poisson GAM. Fitting our model to data collected in seven North Carolina counties from 1999 through 2001, we found that an increase in PM(2.5) exposure is linked to increased risk of cardiovascular mortality in the same day and next 2 days. Specifically, a 10- microg/m3 increase in average PM(2.5) exposure is associated with a 2.5% increase in the relative risk of current-day cardiovascular mortality, a 4.0% increase in the relative risk of cardiovascular mortality the next day, and an 11.4% increase in the relative risk of cardiovascular mortality 2 days later. Because of the small sample size of our study, only the third effect was found to have > 95% posterior probability of being > 0. In addition, we compared the results obtained from our model to those obtained by applying frequentist (or classical, repeated sampling-based) and Bayesian versions of the classical Poisson GAM to our study population.

Adolescent↗

Hierarchical polytomous regression models with applications to health services research.

The analysis of variations is an important area of interest in health services and outcomes research and has two main goals: to identify and quantify variability across units, such as geographic regions or health care providers, in terms of procedure utilization and outcomes, and to explore the links between process, such as regional or hospital practice patterns, and outcomes, such as patient mortality and functional status. Hierarchical regression models are well suited for this type of analysis. In this paper we formulate a hierarchical polytomous regression model and apply it to the analysis of variations in the utilization of alternative cardiac procedures in a national cohort of elderly Medicare patients who had an acute myocardial infarction during 1987. The model is designed to accommodate clustered multinomial data with covariate vectors available on individual cases and on clusters. We present a Bayesian approach to fitting and checking the model using simulated values from the posterior distribution of the parameters. The simulation algorithms are based on Gibbs sampling in combination with Metropolis steps. Using the hierarchical polytomous regression model, we examine how the rates of cardiac procedures depend on patient-level characteristics, including age, gender and race, and whether there exist interstate differences and regional patterns in the use of these procedures.

Aged↗

Using Bayesian statistics to estimate the coefficients of a two- component second-order chlorine bulk decay model for a water distribution system.

Most chlorine decay models for the bulk phase in a water distribution system consider only chlorine concentration and time. Clark [1998. Chlorine demand and trihalomethane formation kinetics: a second-order model. J. Environ. Eng. 124(1), 16-24] first proposed a two-component second-order chlorine decay model based on the concept of competing reacting substances. A corrected mathematical formulation is developed and, because the recent findings suggested that not all natural organic matter (NOM) is involved in the chlorine decay process, an additional parameter is introduced. A parameter assignment method employing Bayesian statistical analysis incorporating Monte Carlo Markov chain (MCMC) with Gibbs sampling to make inferences, is employed in the estimation of model parameters. Three parameters are estimated for the model, namely the ratio of chlorine to TOC, the chlorine reaction rate, and a fraction factor of TOC which represents the true amount of TOC involved in chlorine decay process. Water samples taken from Goderich in the summer of 2005, are used for estimating the parameters.

Bayes Theorem↗

CASPAR: a hierarchical bayesian approach to predict survival times in cancer from gene expression data.

MOTIVATION: DNA microarrays allow the simultaneous measurement of thousands of gene expression levels in any given patient sample. Gene expression data have been shown to correlate with survival in several cancers, however, analysis of the data is difficult, since typically at most a few hundred patients are available, resulting in severely underdetermined regression or classification models. Several approaches exist to classify patients in different risk classes, however, relatively little has been done with respect to the prediction of actual survival times. We introduce CASPAR, a novel method to predict true survival times for the individual patient based on microarray measurements. CASPAR is based on a multivariate Cox regression model that is embedded in a Bayesian framework. A hierarchical prior distribution on the regression parameters is specifically designed to deal with high dimensionality (large number of genes) and low sample size settings, that are typical for microarray measurements. This enables CASPAR to automatically select small, most informative subsets of genes for prediction. RESULTS: Validity of the method is demonstrated on two publicly available datasets on diffuse large B-cell lymphoma (DLBCL) and on adenocarcinoma of the lung. The method successfully identifies long and short survivors, with high sensitivity and specificity. We compare our method with two alternative methods from the literature, demonstrating superior results of our approach. In addition, we show that CASPAR can further refine predictions made using clinical scoring systems such as the International Prognostic Index (IPI) for DLBCL and clinical staging for lung cancer, thus providing an additional tool for the clinician. An analysis of the genes identified confirms previously published results, and furthermore, new candidate genes correlated with survival are identified.

Bayes Theorem↗

Inferring parameters shaping amino acid usage in prokaryotic genomes via Bayesian MCMC methods.

Molar content of guanine plus cytosine (G + C) and optimal growth temperature (OGT) are main factors characterizing the frequency distribution of amino acids in prokaryotes. Previous work, using multivariate exploratory methods, has emphasized ascertainment of biological factors underlying variability between genomes, but the strength of each identified factor on amino acid content has not been quantified. We combine the flexibility of the phylogenetic mixed model (PMM) with the power of Bayesian inference via Markov Chain Monte Carlo (MCMC) methods, to obtain a novel evolutionary picture of amino acid usage in prokaryotic genomes. We implement a Bayesian PMM which incorporates the feature that evolutionary history makes observed data interdependent. As in previous studies with PMM, we present a variance partition; however, attention is also given to the posterior distribution of "systematic effects" that may shed light about the relative importance of and relationships between evolutionary forces acting at the genomic level. In particular, we analyzed influences of G + C, OGT, and respiratory metabolism. Estimates of G + C effects were significant for amino acids coded by G + C or molar content of adenine plus thymine (A + T) in first and second bases. OGT had an important effect on 12 amino acids, probably reflecting complex patterns of protein modifications, to cope with varying environments. The effect of respiratory metabolism was less clear, probably due to the already reported association of G + C with aerobic metabolism. A "heritability" parameter was always high and significant, reinforcing the importance of accommodating phylogenetic relationships in these analyses. "Heritable" component correlations displayed a pattern that tended to cluster "pure" G + C (A + T) in first and second codon positions, suggesting an inherited departure from linear regression on G + C.

Amino Acids↗

Bayesian methods and optimal experimental design for gene mapping by radiation hybrids.

Radiation hybrid mapping is a somatic cell technique for ordering human loci along a chromosome and estimating the physical distance between adjacent loci. The present paper considers a realistic model of fragment generation and retention. This model assumes that fragments are generated in the ancestral cell of a clone according to a Poisson breakage process along the chromosome. Once generated, fragments are independently retained in the clone with a common retention probability. Based on this and less restrictive models, statistical criteria such as minimum obligate breaks, maximum likelihood, and Bayesian posterior probabilities can be used to decide order. Distances can be estimated by either maximum likelihood or Bayesian posterior means. The model also permits rational design of radiation dose for optimal statistical precision. A brief examination of some real data illustrates our criteria and computational algorithms.

Algorithms↗

A Bayesian analysis of amalgam restorations in the Royal Air Force using the counting process approach with nested frailty effects.

Survival analysis methods are increasingly used in dental research to measure risk of tooth eruption and caries as well as life spans of amalgam restorations. Analyses have been extended to account for lack of independence in the data, which arises from the clustering of observations within units such as tooth-surfaces, teeth and subjects. There are various analytical strategies and modelling approaches now available to us in dealing with clustered dental data. In this article, the modelling strategy of Cox's proportional hazards regression is formulated using the counting process approach, which can easily be extended to include time-variant covariates as well as nested random frailty effects. A semi-parametric Bayesian method is presented for the analysis of the proposed model. The methodology is applied to an analysis of nested clustered data on life-span of amalgam restorations in the UK Royal Air Force. These data have previously been analysed using a non-Bayesian approach. The Gibbs sampler, a Markov chain Monte Carlo method, is used to generate samples from the marginal posterior distribution of the parameters of this Bayesian model.

Bayes Theorem↗

The analysis of ring-recovery data using random effects.

We show how random terms, describing both yearly variation and overdispersion, can easily be incorporated into models for mark-recovery data, through the use of Bayesian methods. For recovery data on lapwings, we show that the incorporation of the random terms greatly improves the goodness of fit. Omitting the random terms can lead to overestimation of the significance of weather on survival, and overoptimistic prediction intervals in simulations of future population behavior. Random effects models provide a natural way of modeling overdispersion-which is more satisfactory than the standard classical approach of scaling up all standard errors by a uniform inflation factor. We compare models by means of Bayesian p-values and the deviance information criterion (DIC).

Animals↗

A close relationship between Cercozoa and Foraminifera supported by phylogenetic analyses based on combined amino acid sequences of three cytoskeletal proteins (actin, alpha-tubulin, and beta-tubulin).

Recently, there has been increasing molecular evidence of phylogenetic affinity between Cercozoa and Foraminifera in the eukaryotic lineage. We performed phylogenetic analyses based on the combined (concatenated) amino acid sequence data of actin, alpha-tubulin, and beta-tubulin from a wide variety of eukaryotes, including the foraminifers Planoglabratella opercularis and Reticulomyxa filosa, as well as cercomonad and chlorarachniophyte members of Cercozoa. A monophyletic lineage composed of two foraminiferan species branched with the centroheliozoan species Raphidiophrys contractilis was reconstructed in both Bayesian and maximum-likelihood (ML) analyses under 'linked' models, enforcing a single set of the parameters (the parameter for among-site rate variation and branch lengths) on the entire combined alignment. Considering the extremely divergent nature of Foraminifera and Raphidiophyrs tubulins, the union of these lineages recovered is most probably a long-branch attraction artifact due to ignoring gene-specific evolutionary processes. On the other hand, the foraminiferan lineage was within the radiation of Cercozoa in Bayesian analyses under 'unlinked' model conditions, accommodating differences in evolutionary processes across the three genes in the combined alignment. The Foraminifera+Cercozoa affinity recovered in the latter multi-gene analyses is most likely genuine, and thus our data presented here provide further support for the close relationship between these two protist lineages.

Actins↗

Dynamic survival models with spatial frailty.

In many survival studies, covariates effects are time-varying and there is presence of spatial effects. Dynamic models can be used to cope with the variations of the effects and spatial components are introduced to handle spatial variation. This paper proposes a methodology to simultaneously introduce these components into the model. A number of specifications for the spatial components are considered. Estimation is performed via a Bayesian approach through Markov chain Monte Carlo methods. Models are compared to assess relevance of their components. Analysis of a real data set is performed, showing the relevance of both time-varying covariate effects and spatial components. Extensions to the methodology are proposed along with concluding remarks.

Bayes Theorem↗

Genetic change for clinical mastitis in Norwegian cattle: a threshold model analysis.

Records of clinical mastitis on 1.6 million first-lactation daughters of 2,411 Norwegian Cattle sires that were progeny tested from 1978 through 1998 were analyzed with a threshold model. The main objective was to infer genetic change for the disease in the population. A Bayesian approach via Gibbs sampling was used. The model for the underlying liability had age at first calving, month x year of calving, herd x 3-year-period, and sire of the cow as explanatory variables. Posterior mean (SD) of heritability of liability to clinical mastitis was 0.066 (0.003). Genetic evaluations (posterior means) of sires both in the liability and observable scales were computed. Annual genetic change of liability to clinical mastitis for progeny tested bulls born from 1973 to 1993 was assessed. The linear regression of mean sire effect on year of birth had a posterior mean (SD) of -0.00018 (0.0004), suggesting a nearly constant genetic level for clinical mastitis. However, an analysis of sire posterior means by birth-year of daughters indicated an approximately constant genetic level in the cow population from 1976 to 1990 (-0.02%/yr), and a genetic improvement thereafter (-0.27%/yr). This reflects more emphasis on mastitis in selection of bulls in recent years. Corresponding results obtained with a standard linear model analysis were -0.01% and -0.23% per year, respectively (regression of sire predicted transmitting ability on birth-year of daughters). Genetic change seems to be slightly understated with the linear model, assuming the threshold model holds true.

Animals↗

Bayesian inference on genetic merit under uncertain paternity.

A hierarchical animal model was developed for inference on genetic merit of livestock with uncertain paternity. Fully conditional posterior distributions for fixed and genetic effects, variance components, sire assignments and their probabilities are derived to facilitate a Bayesian inference strategy using MCMC methods. We compared this model to a model based on the Henderson average numerator relationship (ANRM) in a simulation study with 10 replicated datasets generated for each of two traits. Trait 1 had a medium heritability (h2) for each of direct and maternal genetic effects whereas Trait 2 had a high h2 attributable only to direct effects. The average posterior probabilities inferred on the true sire were between 1 and 10% larger than the corresponding priors (the inverse of the number of candidate sires in a mating pasture) for Trait 1 and between 4 and 13% larger than the corresponding priors for Trait 2. The predicted additive and maternal genetic effects were very similar using both models; however, model choice criteria (Pseudo Bayes Factor and Deviance Information Criterion) decisively favored the proposed hierarchical model over the ANRM model.

Animals↗

Random regression models for male and female fertility evaluation using longitudinal binary data.

A longitudinal Bayesian threshold analysis of insemination outcomes was carried out using 2 random regression models with 3 (Model 1) and 5 (Model 2) parameters to model the additive genetic values at the liability scale. All insemination events of first-parity Holstein cows were used. The outcome of an insemination event was treated as a binary response of either a success (1) or a failure (0). Thus, all breeding information for a cow, including all service sires, was included, thereby allowing for a joint evaluation of male and female fertility. An edited data set of 369,353 insemination records from 210,373 first-lactation cows was used. On the liability scale, both models included the systematic effects of herd-year, month of insemination, technician, and regressions on age of service sire and milk yield during the first 100 d of lactation. The random effects in the model were the 3 or 5 random regression coefficients specific to each cow, the permanent effect of the cow, and the service sire effect. Using Model 1, the estimated heritability of an insemination outcome decreased from 0.035 at d 50 to 0.032 at d 140 and then increased continuously with DIM. The genetic correlations for insemination success at different time points ranged from 0.83 to 0.99, and their magnitude decreased with an increase in the interval between inseminations. A similar trend was observed for heritability and genetic correlations using Model 2. However, the average estimate of heritability was much higher (0.058) than those obtained using Model 1 or a repeatability model. In addition, the estimated genetic correlations followed the same trend as Model 1, but were lower and with a higher rate of decrease when the interval between inseminations increased. The posterior mean of service sire variance was 0.01 for both models, and permanent environmental variance was 0.05 and 0.02 for Models 1 and 2, respectively. Model comparison based on the Bayes factor indicated that Model 1 was more plausible, given the data.

Animals↗

An extension of the Cormack-Jolly-Seber model for continuous covariates with application to Microtus pennsylvanicus.

Recent developments in the Cormack-Jolly-Seber (CJS) model for analyzing capture-recapture data have focused on allowing the capture and survival rates to vary between individuals. Several methods have been developed in which capture and survival are functions of auxiliary variables that may be discrete, constant over time, or apply to the population as a whole, but the problem has not been solved for continuous covariates that vary with both time and individual. This article proposes a new method to handle such covariates by modeling changes over time via a diffusion process and using logistic functions to link the variable to the CJS capture and survival rates. Bayesian methods are used to estimate the model parameters. The method is applied to study the effect of body mass on the survival of the North American meadow vole, Microtus pennsylvanicus.

Animals↗