Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

A hierarchical model for estimating response time distributions.

We present a statistical model for inference with response time (RT) distributions. The model has the following features. First, it provides a means of estimating the shape, scale, and location (shift) of RT distributions. Second, it is hierarchical and models between-subjects and within-subjects variability simultaneously. Third, inference with the model is Bayesian and provides a principled and efficient means of pooling information across disparate data from different individuals. Because the model efficiently pools information across individuals, it is particularly well suited for those common cases in which the researcher collects a limited number of observations from several participants. Monte Carlo simulations reveal that the hierarchical Bayesian model provides more accurate estimates than several popular competitors do. We illustrate the model by providing an analysis of the symbolic distance effect in which participants can more quickly ascertain the relationship between nonadjacent digits than that between adjacent digits.

Cognition↗

Bayesian implementation of a genetic model-free approach to the meta-analysis of genetic association studies.

A genetic model-free method for the meta-analysis of genetic association studies is described that estimates the mode of inheritance from the data rather than assuming that it is known. For a bi-allelic polymorphism, with G as risk allele and g as wild-type, the genetic model depends on the ratio of the two log odds ratios, lambda = log OR(Gg)/log OR(GG), where OR(GG) compares GG with gg and OR(Gg) compares Gg with gg. Modelling log OR(GG) as a random effect creates a hierarchical model that can be implemented within a Bayesian framework. In Bayesian modelling, vague prior distributions have to be specified for all unknown parameters when no external information is available. When the data are sparse even supposedly vague prior distributions may have an influence on the posterior estimates. We investigate the impact of different vague prior distributions for the between-study standard deviation of log OR(GG) and for lambda, by considering three published meta-analyses and associated simulations. Our results show that depending on the characteristics of the meta-analysis the results may indeed be sensitive to the choice of vague prior distribution for either parameter. Genetic association studies usually use a case-control design that should be analysed by the corresponding retrospective likelihood. However, under some circumstances the prospective likelihood has been shown to produce identical results and it is usually preferred for its simplicity. In our meta-analyses the two likelihoods give very similar results.

Bayes Theorem↗

A two-component model for counts of infectious diseases.

We propose a stochastic model for the analysis of time series of disease counts as collected in typical surveillance systems on notifiable infectious diseases. The model is based on a Poisson or negative binomial observation model with two components: a parameter-driven component relates the disease incidence to latent parameters describing endemic seasonal patterns, which are typical for infectious disease surveillance data. An observation-driven or epidemic component is modeled with an autoregression on the number of cases at the previous time points. The autoregressive parameter is allowed to change over time according to a Bayesian changepoint model with unknown number of changepoints. Parameter estimates are obtained through the Bayesian model averaging using Markov chain Monte Carlo techniques. We illustrate our approach through analysis of simulated data and real notification data obtained from the German infectious disease surveillance system, administered by the Robert Koch Institute in Berlin. Software to fit the proposed model can be obtained from http://www.statistik.lmu.de/ approximately mhofmann/twins.

Algorithms↗

Search for predictive generic model of aqueous solubility using Bayesian neural nets.

Several predictive models of aqueous solubility have been published. They have good performances on the data sets which have been used for training the models, but usually these data sets do not contain many structures similar to the structures of interest to the drug research and their applicability in drug hunting is questionable. A very diverse data set has been gathered with compounds issued from literature reports and proprietary compounds. These compounds have been grouped in a so-called literature data set I, a proprietary data set II, and a mixed data set III formed by I and II. About 100 descriptors emphasizing surface properties were calculated for every compound. Bayesian learning of neural nets which cumulates the advantages of neural nets without having their weaknesses was used to select the most parsimonious models and train them, from I, II, and III. The models were established by either selecting the most efficient descriptors one by one using a modified Gram-Schmidt procedure (GS) or by simplifying a most complete model using automatic relevance procedure (ARD). The predictive ability of the models was accessed using validation data sets as much unrelated to the training sets as possible, using two new parameters: NDD(x,ref) the normalized smallest descriptor distance of a compound x to a reference data set and CD(x,mod) the combination of NDD(x,ref) with the dispersion of the Bayesian neural nets calculations. The results show that it is possible to obtain a generic predictive model from database I but that the diversity of database II is too restricted to give a model with good generalization ability and that the ARD method applied to the mixed database III gives the best predictive model.

Journal Article↗

Risk assessment of dietary exposure to pesticides using a Bayesian method.

Risk assessment of pesticides can be a statistically difficult problem because pesticides occur only occasionally, but they may occur on multiple components in the diet. A Bayesian statistical model is presented which incorporates multivariate modelling of food consumption and modelling of pesticide measurements which are for a large part below a measurement threshold. It is shown that Bayesian modelling is feasible for a limited number of food components, and that in a data-rich situation the model compares well with an empirical Monte Carlo modelling.

Bayes Theorem↗

Tumour matrilysin expression predicts metastatic potential of stage I (pT1) colon and rectal cancers.

BACKGROUND AND AIMS: Nodal metastases are indisputable determinants of prognosis for colon and rectal cancer. Using classical histological criteria, many attempts to predict nodal metastasis have failed, preventing the adequate management of stage I (pT1) cancer. We investigated the role of tumour matrilysin in predicting metastatic potential, and discuss its potential use in individualising treatment of pT1 colon and rectal cancer. METHODS: The gene signature associated with nodal metastasis was investigated by cDNA array in 24 colon and rectal cancers. We studied 494 colon and rectal cancer patients to identify risk factors for nodal metastasis and evaluated the potential to predict nodal metastasis by either the logistic regression model or the Bayesian neural network model with built-in matrilysin. We then inferred possible causality of nodal metastasis from structural equation modelling. RESULTS: cDNA array revealed that matrilysin was maximally upregulated in the metastasis signature identified. Tumour matrilysin expression emerged as a stage independent risk factor for nodal metastasis, resulting in a similar predictive performance in receiver operating characteristic curve analysis in the two models. A Bayesian approach called automatic relevance determination identified matrilysin as one of the most relevant predictors examined. Structural equation modelling suggested possible direct causality between matrilysin and nodal metastasis. CONCLUSIONS: We have provided evidence that tumour matrilysin expression is a promising biomarker predicting nodal metastasis of colon and rectal cancer. Analysis of tumour matrilysin expression would help clinicians achieve the goal of individualised cancer treatment based on the metastatic potential of pT1 colon and rectal cancer.

Adenocarcinoma↗

Bayesian information criterion for censored survival models.

We investigate the Bayesian Information Criterion (BIC) for variable selection in models for censored survival data. Kass and Wasserman (1995, Journal of the American Statistical Association 90, 928-934) showed that BIC provides a close approximation to the Bayes factor when a unit-information prior on the parameter space is used. We propose a revision of the penalty term in BIC so that it is defined in terms of the number of uncensored events instead of the number of observations. For a simple censored data model, this revision results in a better approximation to the exact Bayes factor based on a conjugate unit-information prior. In the Cox proportional hazards regression model, we propose defining BIC in terms of the maximized partial likelihood. Using the number of deaths rather than the number of individuals in the BIC penalty term corresponds to a more realistic prior on the parameter space and is shown to improve predictive performance for assessing stroke risk in the Cardiovascular Health Study.

Aged↗

A class of Bayesian shared gamma frailty models with multivariate failure time data.

For multivariate failure time data, we propose a new class of shared gamma frailty models by imposing the Box-Cox transformation on the hazard function, and the product of the baseline hazard and the frailty. This novel class of models allows for a very broad range of shapes and relationships between the hazard and baseline hazard functions. It includes the well-known Cox gamma frailty model and a new additive gamma frailty model as two special cases. Due to the nonnegative hazard constraint, this shared gamma frailty model is computationally challenging in the Bayesian paradigm. The joint priors are constructed through a conditional-marginal specification, in which the conditional distribution is univariate, and it absorbs the nonlinear parameter constraints. The marginal part of the prior specification is free of constraints. The prior distributions allow us to easily compute the full conditionals needed for Gibbs sampling, while incorporating the constraints. This class of shared gamma frailty models is illustrated with a real dataset.

Adolescent↗

Lung cancer rate predictions using generalized additive models.

Predictions of lung cancer incidence and mortality are necessary for planning public health programs and clinical services. It is proposed that generalized additive models (GAMs) are practical for cancer rate prediction. Smooth equivalents for classical age-period, age-cohort, and age-period-cohort models are available using one-dimensional smoothing splines. We also propose using two-dimensional smoothing splines for age and period. Variance estimation can be based on the bootstrap. To assess predictive performance, we compared the models with a Bayesian age-period-cohort model. Model comparison used cross-validation and measures of predictive performance for recent predictions. The models were applied to data from the World Health Organization Mortality Database for females in five countries. Model choice between the age-period-cohort models and the two-dimensional models was equivocal with respect to cross-validation, while the two-dimensional GAMs had very good predictive performance. The Bayesian model performed poorly due to imprecise predictions and the assumption of linearity outside of observed data. In summary, the two-dimensional GAM performed well. The GAMs make the important prediction that female lung cancer rates in these countries will be stable or begin to decline in the future.

Adult↗

Modeling cellular processes with variational Bayesian cooperative vector quantizer.

Gene expression of a cell is controlled by sophisticated cellular processes. The capability of inferring the states of these cellular processes would provide insight into the mechanism of gene expression control system. In this paper, we propose and investigate the cooperative vector quantizer (CVQ) model for analysis of microarray data. The CVQ model could be capable of decomposing observed microarray data into many different regulatory subprocesses. To make the CVQ analysis tractable we develop and apply variational approximations. Bayesian model selection is employed in the model, so that the optimal number processes is determined purely from observed micro-array data. We test the model and algorithms on two datasets: (1) simulated gene-expression data and (2) real-world yeast cell-cycle microarray data. The results illustrate the ability of the CVQ approach to recover and characterize regulatory gene expression subprocesses, indicating a potential for advanced gene expression data analysis.

Algorithms↗

Clustering of genes into regulons using integrated modeling-COGRIM.

We present a Bayesian hierarchical model and Gibbs Sampling implementation that integrates gene expression, ChIP binding, and transcription factor motif data in a principled and robust fashion. COGRIM was applied to both unicellular and mammalian organisms under different scenarios of available data. In these applications, we demonstrate the ability to predict gene-transcription factor interactions with reduced numbers of false-positive findings and to make predictions beyond what is obtained when single types of data are considered.

CCAAT-Enhancer-Binding Protein-beta↗

Oligogenic model selection using the Bayesian Information Criterion: linkage analysis of the P300 Cz event-related brain potential.

The traditional likelihood-based approach to hypothesis testing may not be an optimal strategy for evaluating oligogenic models of inheritance. Under oligogenic inheritance the number of possible multilocus models can become very large; there may be several competing linkage models having similar likelihoods; and comparisons among non-nested models can be required to determine if a given multilocus model provides a significantly better fit to observed phenotypic variation than an alternative model. We propose an efficient Bayesian approach to oligogenic model selection that makes use of existing model likelihoods, and show how model uncertainty can be incorporated into parameter estimation.

Alcoholism↗

Bayesian methods for analysis of binary outcome data in cluster randomized trials on the absolute risk scale.

A Bayesian hierarchical modelling approach to the analysis of cluster randomized trials has advantages in terms of allowing for full parameter uncertainty, flexible modelling of covariates and variance structure, and use of prior information. Previously, such modelling of binary outcome data required use of a log-odds ratio scale for the treatment effect estimate and an approximation linking the intracluster correlation (ICC) to the between-cluster variance on a log-odds scale. In this paper we develop this method to allow estimation on the absolute risk scale, which facilitates clinical interpretation of both the treatment effect and the between-cluster variance. We describe a range of models and apply them to data from a trial of different interventions to promote secondary prevention of coronary heart disease in primary care. We demonstrate how these models can be used to incorporate prior data about typical ICCs, to derive a posterior distribution for the number needed to treat, and to consider both cluster and individual level covariates. Using these methods, we can benefit from the advantages of Bayesian modelling of binary outcome data at the same time as providing results on a clinically interpretable scale.

Bayes Theorem↗

A Bayesian threshold-normal mixture model for analysis of a continuous mastitis-related trait.

Mastitis is associated with elevated somatic cell count in milk, inducing a positive correlation between milk somatic cell score (SCS) and the absence or presence of the disease. In most countries, selection against mastitis has focused on selecting parents with genetic evaluations that have low SCS. Univariate or multivariate mixed linear models have been used for statistical description of SCS. However, an observation of SCS can be regarded as drawn from a 2- (or more) component mixture defined by the (usually) unknown health status of a cow at the test-day on which SCS is recorded. A hierarchical 2-component mixture model was developed, assuming that the health status affecting the recorded test-day SCS is completely specified by an underlying liability variable. Based on the observed SCS, inferences can be drawn about disease status and parameters of both SCS and liability to mastitis. The prior probability of putative mastitis was allowed to vary between subgroups (e.g., herds, families), by specifying fixed and random effects affecting both SCS and liability. Using simulation, it was found that a Bayesian model fitted to the data yielded parameter estimates close to their true values. The model provides selection criteria that are more appealing than selection for lower SCS. The proposed model can be extended to handle a wide range of problems related to genetic analyses of mixture traits.

Animals↗

Meta-analysis of sentinel lymph node biopsy after preoperative chemotherapy in patients with breast cancer.

BACKGROUND: Women with breast cancer are more frequently being treated with preoperative neoadjuvant chemotherapy. The reliability of sentinel lymph node biopsy (SLNB) following chemotherapy has not been determined. This was a meta-analysis of studies that examined the results of SLNB after preoperative chemotherapy. METHODS: Included articles had to meet two criteria. First, patients had to have had operable breast cancer and to have undergone SLNB after preoperative chemotherapy and, second, patients had to have undergone subsequent axillary lymph node dissection. Meta-analyses were performed in which Bayesian hierarchical models were created to estimate the identification rate (IR) and sensitivity of SLNB in this setting. RESULTS: Twenty-one studies were identified that included a total of 1273 patients. The IRs reported ranged from 72 to 100 per cent, with a pooled estimate of 90 per cent. The sensitivity of SLNB ranged from 67 to 100 per cent, with a pooled estimate of 88 (95 per cent confidence interval 85 to 90) per cent. Meta-analyses performed using Bayesian modelling resulted in (posterior) estimates for IR and sensitivity of 91 (95 per cent credible interval 88 to 94) and 88 (95 per cent credible interval 84 to 91) per cent respectively. CONCLUSION: SLNB is a reliable tool for planning treatment after preoperative chemotherapy.

Antineoplastic Agents↗

Spatial variation of mortality for common and rare cancers in Piedmont, Italy, from 1980 to 2000: a Bayesian approach.

A Bayesian hierarchical model was used to study the spatial variation in mortality risk from lung and pleural cancer in both sexes, and breast and soft tissue sarcoma (STS) in women in Piedmont (north-west Italy, average population 4 349 411) from 1980 to 2000. Of these four neoplasms, two are common (lung and breast) and two rare (pleura and STS); two have well recognized risk factors (lung and pleura) while the other two (breast and STS) have no single strong risk factor. Data were analysed at a small-area level (1206 municipalities, population 39 to 989 663), using both standardized mortality ratios and Bayesian-estimated mortality risks. The Bayesian model allowed for both heterogeneity (through spatially independent random effects) and clustering (through spatially correlated random effects) and, by borrowing information from neighbouring areas, provided stable estimates for areas with sparse data. The aim was to reduce the noise in the disease maps to highlight the true underlying mortality distribution. Lung cancer in men showed strong spatial structure with a marked east-west gradient, but no appreciable urban-rural differences. In contrast, high mortality areas for female lung cancer were observed around conurbations. Female breast cancer and STS appeared to be spread uniformly across the region. Pleural cancer mortality clusters were evident around areas with major asbestos manufacturers, or natural asbestiform fibre pollution. Maps of Bayesian-estimated mortality risk provided appreciably clearer pictures of risk distribution than did maps of the standardized mortality ratio.

Bayes Theorem↗

Suramin: rapid loading and weekly maintenance regimens for cancer patients.

PURPOSE: Suramin is an anticancer agent with a narrow therapeutic window and a terminal half-life of 45 to 55 days. These characteristics make it necessary to control accurately the serum concentrations of the drug. Therefore, the aim of the present study was to develop a rapid loading regimen, followed by weekly administration of suramin to maintain serum concentrations of between 150 and 300 micrograms/mL for 8 weeks. PATIENTS AND METHODS: Eligible patients were treated with five different loading regimens. Initially, weekly maintenance doses were estimated manually by the treating physician. Subsequently, computer-assisted dosing that used Bayesian pharmacokinetic modeling was used. RESULTS: Thirty-eight courses of suramin that were administered to 35 patients were studied. The optimal loading regimen consisted of a continuous infusion of 600 mg/m2 during a 24-hour period, which resulted in a mean serum concentration of 319 micrograms/mL. Potentially toxic concentrations that were observed with shorter infusions were avoided. Maintenance treatment, which used the weekly administration of suramin during a 6-hour period, seemed to be able to maintain mean suramin serum trough concentrations of 150 micrograms/mL, while preventing mean peak concentrations of more than 300 micrograms/mL. The use of Bayesian pharmacokinetics was superior to manual estimation in tailoring the optimal dose to the therapeutic window. CONCLUSIONS: Continuous infusion is the optimal way of delivering suramin during the loading phase. To maintain trough levels and peak levels within a narrower therapeutic window, suramin will have to be administered more frequently than once a week. Bayesian modeling based on individual serum levels and population pharmacokinetics allows accurate dosing to maintain suramin levels within the therapeutic window.

Adult↗

Practical model-based dose-finding in phase I clinical trials: methods based on toxicity.

We describe two practical, outcome-adaptive statistical methods for dose-finding in phase I clinical trials. One is the continual reassessment method and the other is based on a logistic regression model. Both methods use Bayesian probability models as a basis for learning from the accruing data during the trial, choosing doses for successive patient cohorts, and selecting a maximum tolerable dose (MTD). These methods are illustrated and compared to the conventional 3+3 algorithm by application to a particular trial in renal cell carcinoma. We also compare their average behavior by computer simulation under each of several hypothetical dose-toxicity curves. The comparisons show that the Bayesian methods are much more reliable than the conventional algorithm for selecting an MTD, and that they have a low risk of treating patients at unacceptably toxic doses.

Algorithms↗