Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

A variational Bayesian mixture modelling framework for cluster analysis of gene-expression data.

MOTIVATION: Accurate subcategorization of tumour types through gene-expression profiling requires analytical techniques that estimate the number of categories or clusters rigorously and reliably. Parametric mixture modelling provides a natural setting to address this problem. RESULTS: We compare a criterion for model selection that is derived from a variational Bayesian framework with a popular alternative based on the Bayesian information criterion. Using simulated data, we show that the variational Bayesian method is more accurate in finding the true number of clusters in situations that are relevant to current and future microarray studies. We also compare the two criteria using freely available tumour microarray datasets and show that the variational Bayesian method is more sensitive to capturing biologically relevant structure.

Algorithms↗

Expert systems in psychiatry. A review.

Existing computer-based decision aids in the areas of psychiatric diagnosis and consultation are reviewed, and the prospects for expert system development within the mental health field are discussed. Emphasis is placed upon the decision-making models used in these systems rather than on their particular application area. The decision-making paradigms discussed are (1) data bank analysis, (2) statistical pattern recognition, (3) Bayesian analysis, (4) logical flow chart method, and (5) knowledge-based (expert system) approaches. For each paradigm, its essential features, its strengths and weaknesses, and some example applications are presented.

Diagnosis, Computer-Assisted↗

Inferring gene networks from time series microarray data using dynamic Bayesian networks.

Dynamic Bayesian networks (DBNs) are considered as a promising model for inferring gene networks from time series microarray data. DBNs have overtaken Bayesian networks (BNs) as DBNs can construct cyclic regulations using time delay information. In this paper, a general framework for DBN modelling is outlined. Both discrete and continuous DBN models are constructed systematically and criteria for learning network structures are introduced from a Bayesian statistical viewpoint. This paper reviews the applications of DBNs over the past years. Real data applications for Saccharomyces cerevisiae time series gene expression data are also shown.

Algorithms↗

Flexible dose-response models for Japanese atomic bomb survivor data: Bayesian estimation and prediction of cancer risk.

Generalised absolute risk models were fitted to the latest Japanese atomic bomb survivor cancer incidence data using Bayesian Markov Chain Monte Carlo methods, taking account of random errors in the DS86 dose estimates. The resulting uncertainty distributions in the relative risk model parameters were used to derive uncertainties in population cancer risks for a current UK population. Because of evidence for irregularities in the low-dose dose response, flexible dose-response models were used, consisting of a linear-quadratic-exponential model, used to model the high-dose part of the dose response, together with piecewise-linear adjustments for the two lowest dose groups. Following an assumed administered dose of 0.001 Sv, lifetime leukaemia radiation-induced incidence risks were estimated to be 1.11 x 10(-2) Sv(-1) (95% Bayesian CI -0.61, 2.38) using this model. Following an assumed administered dose of 0.001 Sv, lifetime solid cancer radiation-induced incidence risks were calculated to be 7.28 x 10(-2) Sv(-1) (95% Bayesian CI -10.63, 22.10) using this model. Overall, cancer incidence risks predicted by Bayesian Markov Chain Monte Carlo methods are similar to those derived by classical likelihood-based methods and which form the basis of established estimates of radiation-induced cancer risk.

Algorithms↗

The latent process decomposition of cDNA microarray data sets.

We present a new computational technique (a software implementation, data sets, and supplementary information are available at http://www.enm.bris.ac.uk/lpd/) which enables the probabilistic analysis of cDNA microarray data and we demonstrate its effectiveness in identifying features of biomedical importance. A hierarchical Bayesian model, called Latent Process Decomposition (LPD), is introduced in which each sample in the data set is represented as a combinatorial mixture over a finite set of latent processes, which are expected to correspond to biological processes. Parameters in the model are estimated using efficient variational methods. This type of probabilistic model is most appropriate for the interpretation of measurement data generated by cDNA microarray technology. For determining informative substructure in such data sets, the proposed model has several important advantages over the standard use of dendrograms. First, the ability to objectively assess the optimal number of sample clusters. Second, the ability to represent samples and gene expression levels using a common set of latent variables (dendrograms cluster samples and gene expression values separately which amounts to two distinct reduced space representations). Third, in constrast to standard cluster models, observations are not assigned to a single cluster and, thus, for example, gene expression levels are modeled via combinations of the latent processes identified by the algorithm. We show this new method compares favorably with alternative cluster analysis methods. To illustrate its potential, we apply the proposed technique to several microarray data sets for cancer. For these data sets it successfully decomposes the data into known subtypes and indicates possible further taxonomic subdivision in addition to highlighting, in a wholly unsupervised manner, the importance of certain genes which are known to be medically significant. To illustrate its wider applicability, we also illustrate its performance on a microarray data set for yeast.

Algorithms↗

Using graphical models and genomic expression data to statistically validate models of genetic regulatory networks.

We propose a model-driven approach for analyzing genomic expression data that permits genetic regulatory networks to be represented in a biologically interpretable computational form. Our models permit latent variables capturing unobserved factors, describe arbitrarily complex (more than pair-wise) relationships at varying levels of refinement, and can be scored rigorously against observational data. The models that we use are based on Bayesian networks and their extensions. As a demonstration of this approach, we utilize 52 genomes worth of Affymetrix GeneChip expression data to correctly differentiate between alternative hypotheses of the galactose regulatory network in S. cerevisiae. When we extend the graph semantics to permit annotated edges, we are able to score models describing relationships at a finer degree of specification.

Bayes Theorem↗

Bayesian modelling of multivariate quantitative traits using seemingly unrelated regressions.

We investigate a Bayesian approach to modelling the statistical association between markers at multiple loci and multivariate quantitative traits. In particular, we describe the use of Bayesian Seemingly Unrelated Regressions (SUR) whereby genotypes at the different loci are allowed to have non-simultaneous effects on the phenotypes considered with residuals from each regression assumed correlated. We present results from simulations showing that, under rather general conditions that are likely to hold in real situations, the Bayesian SUR approach has increased probability of selecting the true model compared to univariate analyses. Finally, we apply our methods to data from subjects genotyped for 12 SNPs in the apolipoprotein E (APOE) gene. Phenotypes relate to response to treatment with atorvastatin and include changes in total cholesterol, low-density lipoprotein cholesterol, and triglycerides. Missing genotype data are naturally accommodated in our Bayesian framework by imputing them using a nested haplotype phasing algorithm.

Algorithms↗

Variational Bayesian inference for fMRI time series.

We describe a Bayesian estimation and inference procedure for fMRI time series based on the use of General Linear Models with Autoregressive (AR) error processes. We make use of the Variational Bayesian (VB) framework which approximates the true posterior density with a factorised density. The fidelity of this approximation is verified via Gibbs sampling. The VB approach provides a natural extension to previous Bayesian analyses which have used Empirical Bayes. VB has the advantage of taking into account the variability of hyperparameter estimates with little additional computational effort. Further, VB allows for automatic selection of the order of the AR process. Results are shown on simulated data and on data from an event-related fMRI experiment.

Algorithms↗

Estimating vaccine efficacy using auxiliary outcome data and a small validation sample.

In vaccine studies, a specific diagnosis of a suspected case by culture or serology of the infectious agent is expensive and difficult. Implementing validation sets in the study is less expensive and is easier to carry out. In studies using validation sets, the non-specific or auxiliary outcome is measured on each participant while the specific outcome is measured only for a small proportion of the participants. Vaccine efficacy, defined as one minus some measure of relative risk, could be severely attenuated if based only on the auxiliary outcome. Applying missing data analysis techniques could thus correct the bias while maintaining statistical efficiency. However, when the sample size in the validation sets is small and the vaccine is highly efficacious, all specific outcomes are likely to be negative in the validation set in the vaccinated group. Two commonly used missing data analysis methods, the mean score method and multiple imputation, depend on the ad hoc continuity correction when none of the specific outcomes are positive and the normality or log-normality assumption of relative risk, which may not hold when the relative risk is highly skewed, to estimate the confidence interval. In this paper, we propose a Bayesian method to estimate vaccine efficacy and its highest probability density (HPD) credible set using Monte Carlo (MC) methods when using auxiliary outcome data and a small validation sample. Comparing the performance of these approaches using data from a field study of influenza vaccine and simulations, we recommend to use the Bayesian method in this situation.

Adolescent↗

Bayesian modelling of extreme surges on the UK east coast.

The catastrophic surge event of 1953 on the eastern UK and northern European coastlines led to widespread agreement on the necessity of a coordinated response to understand the risk of future oceanographic flood events and, so far as possible, to afford protection against such events. One element of this response was better use of historical data and scientific knowledge in assessing flood risk. The timing of the event also coincided roughly with the birth of extreme value theory as a statistical discipline for measuring risks of extreme events, and over the last 50 years, as techniques have been developed and refined, various attempts have been made to improve the precision of flood risk assessment around the UK coastline. In part, this article provides a review of such developments. Our broader aim, however, is to show how modern statistical modelling techniques, allied with the tools of extreme value theory and knowledge of sea-dynamic physics, can lead to further improvements in flood risk assessment. Our long-term goal is a coherent spatial model that exploits spatial smoothness in the surge process characteristics and we outline the details of such a model. The analysis of the present article, however, is restricted to a site-by-site analysis of high-tide surges. Nonetheless, we argue that the Bayesian methodology adopted for such analysis enables a risk-based interpretation of results that is most natural in this setting, and preferable to inferences that are available from more conventional analyses.

Bayes Theorem↗

Parametric and nonparametric population methods: their comparative performance in analysing a clinical dataset and two Monte Carlo simulation studies.

BACKGROUND AND OBJECTIVES: This study examined parametric and nonparametric population modelling methods in three different analyses. The first analysis was of a real, although small, clinical dataset from 17 patients receiving intramuscular amikacin. The second analysis was of a Monte Carlo simulation study in which the populations ranged from 25 to 800 subjects, the model parameter distributions were Gaussian and all the simulated parameter values of the subjects were exactly known prior to the analysis. The third analysis was again of a Monte Carlo study in which the exactly known population sample consisted of a unimodal Gaussian distribution for the apparent volume of distribution (V(d)), but a bimodal distribution for the elimination rate constant (k(e)), simulating rapid and slow eliminators of a drug. METHODS: For the clinical dataset, the parametric iterative two-stage Bayesian (IT2B) approach, with the first-order conditional estimation (FOCE) approximation calculation of the conditional likelihoods, was used together with the nonparametric expectation-maximisation (NPEM) and nonparametric adaptive grid (NPAG) approaches, both of which use exact computations of the likelihood. For the first Monte Carlo simulation study, these programs were also used. A one-compartment model with unimodal Gaussian parameters V(d) and k(e) was employed, with a simulated intravenous bolus dose and two simulated serum concentrations per subject. In addition, a newer parametric expectation-maximisation (PEM) program with a Faure low discrepancy computation of the conditional likelihoods, as well as nonlinear mixed-effects modelling software (NONMEM), both the first-order (FO) and the FOCE versions, were used. For the second Monte Carlo study, a one-compartment model with an intravenous bolus dose was again used, with five simulated serum samples obtained from early to late after dosing. A unimodal distribution for V(d) and a bimodal distribution for k(e) were chosen to simulate two subpopulations of 'fast' and 'slow' metabolisers of a drug. NPEM results were compared with that of a unimodal parametric joint density having the true population parameter means and covariance. RESULTS: For the clinical dataset, the interindividual parameter percent coefficients of variation (CV%) were smallest with IT2B, suggesting less diversity in the population parameter distributions. However, the exact likelihood of the results was also smaller with IT2B, and was 14 logs greater with NPEM and NPAG, both of which found a greater and more likely diversity in the population studied. For the first Monte Carlo dataset, NPAG and PEM, both using accurate likelihood computations, showed statistical consistency. Consistency means that the more subjects studied, the closer the estimated parameter values approach the true values. NONMEM FOCE and NONMEM FO, as well as the IT2B FOCE methods, do not have this guarantee. Results obtained by IT2B FOCE, for example, often strayed visibly away from the true values as more subjects were studied. Furthermore, with respect to statistical efficiency (precision of parameter estimates), NPAG and PEM had good efficiency and precise parameter estimates, while precision suffered with NONMEM FOCE and IT2B FOCE, and severely so with NONMEM FO. For the second Monte Carlo dataset, NPEM closely approximated the true bimodal population joint density, while an exact parametric representation of an assumed joint unimodal density having the true population means, standard deviations and correlation gave a totally different picture. CONCLUSIONS: The smaller population interindividual CV% estimates with IT2B on the clinical dataset are probably the result of assuming Gaussian parameter distributions and/or of using the FOCE approximation. NPEM and NPAG, having no constraints on the shape of the population parameter distributions, and which compute the likelihood exactly and estimate parameter values with greater precision, detected the more likely greater diversity in the parameter values in the population studied. In the first Monte Carlo study, NPAG and PEM had more precise parameter estimates than either IT2B FOCE or NONMEM FOCE, as well as much more precise estimates than NONMEM FO. In the second Monte Carlo study, NPEM easily detected the bimodal parameter distribution at this initial step without requiring any further information. Population modelling methods using exact or accurate computations have more precise parameter estimation, better stochastic convergence properties and are, very importantly, statistically consistent. Nonparametric methods are better than parametric methods at analysing populations having unanticipated non-Gaussian or multimodal parameter distributions.

Aged↗

Acute middle ear infection in small children: a Bayesian analysis using multiple time scales.

The study is based on a sample of 965 children living in Oulu region (Finland), who were monitored for acute middle ear infections from birth to the age of two years. We introduce a nonparametrically defined intensity model for ear infections, which involves both fixed and time dependent covariates, such as calendar time, current age, length of breast-feeding time until present, or current type of day care. Unmeasured heterogeneity, which manifests itself in frequent infections in some children and rare in others and which cannot be explained in terms of the known covariates, is modelled by using individual frailty parameters. A Bayesian approach is proposed to solve the inferential problem. The numerical work is carried out by Monte Carlo integration (Metropolis-Hastings algorithm).

Acute Disease↗

Comparison of Bayesian and maximum-likelihood inference of population genetic parameters.

UNLABELLED: Comparison of the performance and accuracy of different inference methods, such as maximum likelihood (ML) and Bayesian inference, is difficult because the inference methods are implemented in different programs, often written by different authors. Both methods were implemented in the program MIGRATE, that estimates population genetic parameters, such as population sizes and migration rates, using coalescence theory. Both inference methods use the same Markov chain Monte Carlo algorithm and differ from each other in only two aspects: parameter proposal distribution and maximization of the likelihood function. Using simulated datasets, the Bayesian method generally fares better than the ML approach in accuracy and coverage, although for some values the two approaches are equal in performance. MOTIVATION: The Markov chain Monte Carlo-based ML framework can fail on sparse data and can deliver non-conservative support intervals. A Bayesian framework with appropriate prior distribution is able to remedy some of these problems. RESULTS: The program MIGRATE was extended to allow not only for ML(-) maximum likelihood estimation of population genetics parameters but also for using a Bayesian framework. Comparisons between the Bayesian approach and the ML approach are facilitated because both modes estimate the same parameters under the same population model and assumptions.

Bayes Theorem↗

Probe-level measurement error improves accuracy in detecting differential gene expression.

MOTIVATION: Finding differentially expressed genes is a fundamental objective of a microarray experiment. Numerous methods have been proposed to perform this task. Existing methods are based on point estimates of gene expression level obtained from each microarray experiment. This approach discards potentially useful information about measurement error that can be obtained from an appropriate probe-level analysis. Probabilistic probe-level models can be used to measure gene expression and also provide a level of uncertainty in this measurement. This probe-level measurement error provides useful information which can help in the identification of differentially expressed genes. RESULTS: We propose a Bayesian method to include probe-level measurement error into the detection of differentially expressed genes from replicated experiments. A variational approximation is used for efficient parameter estimation. We compare this approximation with MAP and MCMC parameter estimation in terms of computational efficiency and accuracy. The method is used to calculate the probability of positive log-ratio (PPLR) of expression levels between conditions. Using the measurements from a recently developed Affymetrix probe-level model, multi-mgMOS, we test PPLR on a spike-in dataset and a mouse time-course dataset. Results show that the inclusion of probe-level measurement error improves accuracy in detecting differential gene expression. AVAILABILITY: The MAP approximation and variational inference described in this paper have been implemented in an R package pplr. The MCMC method is implemented in Matlab. Both software are available from http://umber.sbs.man.ac.uk/resources/puma.

Algorithms↗

Constructing molecular classifiers for the accurate prognosis of lung adenocarcinoma.

PURPOSE: Individualized therapy of lung adenocarcinoma depends on the accurate classification of patients into subgroups of poor and good prognosis, which reflects a different probability of disease recurrence and survival following therapy. However, it is currently impossible to reliably identify specific high-risk patients. Here, we propose a computational model system which accurately predicts the clinical outcome of individual patients based on their gene expression profiles. EXPERIMENTAL DESIGN: Gene signatures were selected using feature selection algorithms random forests, correlation-based feature selection, and gain ratio attribute selection. Prediction models were built using random committee and Bayesian belief networks. The prognostic power of the survival predictors was also evaluated using hierarchical cluster analysis and Kaplan-Meier analysis. RESULTS: The predictive accuracy of an identified 37-gene survival signature is 0.96 as measured by the area under the time-dependent receiver operating curves. The cluster analysis, using the 37-gene signature, aggregates the patient samples into three groups with distinct prognoses (Kaplan-Meier analysis, P < 0.0005, log-rank test). All patients in cluster 1 were in stage I, with N0 lymph node status (no metastasis) and smaller tumor size (T1 or T2). Additionally, a 12-gene signature correctly predicts the stage of 94.2% of patients. CONCLUSIONS: Our results show that the prediction models based on the expression levels of a small number of marker genes could accurately predict patient outcome for individualized therapy of lung adenocarcinoma. Such an individualized treatment may significantly increase survival due to the optimization of treatment procedures and improve lung cancer survival every year through the 5-year checkpoint.

Adenocarcinoma↗

Numerical non-identifiability regions of the minimal model of glucose kinetics: superiority of Bayesian estimation.

The so-called minimal model (MM) of glucose kinetics is widely employed to estimate insulin sensitivity (S(I)) both in clinical and epidemiological studies. Usually, MM is numerically identified by resorting to Fisherian parameter estimation techniques, such as maximum likelihood (ML). However, unsatisfactory parameter estimates are sometimes obtained, e.g. S(I) estimates virtually zero or unrealistically high and affected by very large uncertainty, making the practical use of MM difficult. The first result of this paper concerns the mathematical demonstration that these estimation difficulties are inherent to MM structure which can expose S(I) estimation to the risk of numerical non-identifiability. The second result is based on simulation studies and shows that Bayesian parameter estimation techniques are less sensitive, in terms of both accuracy and precision, than the Fisherian ones with respect to these difficulties. In conclusion, Bayesian parameter estimation can successfully deal with difficulties of MM identification inherently due to its structure.

Bayes Theorem↗

An artificial neural network-based expert system for the appraisal of two-car crash accidents.

This paper employs artificial neural network (ANN) to develop an accident appraisal expert system. Two ANN models-- party-based and case-based-- with different hidden neurons are trained and validated by k-fold (k=3) cross validation method. A total of 537 two-car crash accidents (1074 parties involved) are randomly and equally divided into three subsets. For the comparison, a discrimination analysis (DA) model is also calibrated. The results show that the ANN model can achieve a high correctness rate of 85.72% in training and 77.91% in validation and a low Schwarz's Bayesian information criterion (SBC) of -0.82 in training and 0.13 in validation, which indicates that the ANN model is suitable for accident appraisal. Furthermore, in order to measure the importance of each explanatory variable, a general influence (GI) index is computed based on the trained weights of ANN. It is found that the most influential variable is right-of-way, followed by location and alcoholic use. This finding concurs with the prior knowledge in accident appraisal. Thus, for the fair assessment of accident liabilities the correctness of these three key variables is of critical importance to police investigation reports.

Accidents, Traffic↗

Predicting the course of meningococcal disease outbreaks in closed subpopulations.

A stochastic epidemic model was applied to meningococcal disease outbreaks in defined small populations such as military garrisons and schools. Meningococci are spread primarily by asymptomatic carriers and only a small proportion of those infected develop invasive disease. Bayesian predictions of numbers of invasive cases were developed, based on observed data using a stochastic epidemic model. We used additional data sets to model both disease probability and duration of carriage. Markov chain Monte Carlo sampling techniques were used to compute the full posterior distribution which summarized all information drawn together from multiple sources.

Adult↗