Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Application of likelihood ratio and logistic regression models to landslide susceptibility mapping using GIS.

For landslide susceptibility mapping, this study applied and verified a Bayesian probability model, a likelihood ratio and statistical model, and logistic regression to Janghung, Korea, using a Geographic Information System (GIS). Landslide locations were identified in the study area from interpretation of IRS satellite imagery and field surveys; and a spatial database was constructed from topographic maps, soil type, forest cover, geology and land cover. The factors that influence landslide occurrence, such as slope gradient, slope aspect, and curvature of topography, were calculated from the topographic database. Soil texture, material, drainage, and effective depth were extracted from the soil database, while forest type, diameter, and density were extracted from the forest database. Land cover was classified from Landsat TM satellite imagery using unsupervised classification. The likelihood ratio and logistic regression coefficient were overlaid to determine each factor's rating for landslide susceptibility mapping. Then the landslide susceptibility map was verified and compared with known landslide locations. The logistic regression model had higher prediction accuracy than the likelihood ratio model. The method can be used to reduce hazards associated with landslides and to land cover planning.

Databases, Factual↗

A Bayesian approach to Weibull survival models--application to a cancer clinical trial.

In this paper we outline a class of fully parametric proportional hazards models, in which the baseline hazard is assumed to be a power transform of the time scale, corresponding to assuming that survival times follow a Weibull distribution. Such a class of models allows for the possibility of time varying hazard rates, but assumes a constant hazard ratio. We outline how Bayesian inference proceeds for such a class of models using asymptotic approximations which require only the ability to maximize the joint log posterior density. We apply these models to a clinical trial to assess the efficacy of neutron therapy compared to conventional treatment for patients with tumours of the pelvic region. In this trial there was prior information about the log hazard ratio both in terms of elicited clinical beliefs and the results of previous studies. Finally, we consider a number of extensions to this class of models, in particular the use of alternative baseline functions, and the extension to multi-state data.

Bayes Theorem↗

Understanding mathematical models for breast cancer risk assessment and counseling.

Chemoprevention and prophylactic surgery are effective interventions for lowering breast cancer incidence. However, these approaches are associated with risks of their own. Accurate individualized breast cancer risk assessment is an essential component of the risk/benefit analysis that must take place prior to implementing either of these strategies. Several mathematical models for estimating individual breast cancer risk have been proposed over the last decade. The Gail model is the most generally applicable model; however, it neglects family history information in second-degree relatives, treats pre- and postmenopausal breast cancer the same, and ignores personal histories of lobular neoplasia. The Claus model is a better family history model, but it does not assign any special relevance to histories of bilateral breast cancer or ovarian cancer, and neglects all of the nonfamily history information accounted for by the Gail model. BRCAPRO is a Bayesian family history model that calculates individual breast cancer probabilities based on the probability that a family carries a mutation in one of the BRCA genes. Though its treatment of family history information is more thorough than the other models, it neglects the nonfamily history risk factors accounted for by the Gail model and may not appreciate familial clustering unrelated to BRCA gene mutation. A thorough understanding of the principles of risk analysis and the available mathematical models is essential for anyone wishing to perform intervention counseling. This review describes the basic components of risk analysis, explains how the mathematical models work and compares the strengths and weaknesses of the various models. CancerGene is a software tool for running all of these models. It may be obtained without charge at http://www.swmed.edu/home_pages/cancergene.

Breast Neoplasms↗

Late potential recognition by artificial neural networks.

Ventricular late potentials (LP's) are high-frequency low-amplitude signals obtained from signal-averaged electrocardiograms (ECG's) [SAECG's]. LP's are useful in identifying patients prone to ventricular tachycardia (VT), spontaneous or inducible during electrophysiology testing. A combination of self-organizing and supervised artificial neural network (ANN) models was developed to identify patients with a positive electrophysiology (PEP) test for inducible ventricular tachycardia from patients with a negative electrophysiology (NEP) test using LP's. We have added morphology information of vector magnitude waveform to original set of three time-domain features of LP's, which are total QRS duration (TQRSD), high-frequency low-amplitude signal duration (HFLAD), and root-mean-square voltage (RMSV). Pattern recognition results from an ANN model with this combination feature set are superior to the results from Bayesian classification model based on conventional three time-domain features of SAECG. In order to increase the robustness of the recognition, a filtered QRS offset point is randomly shifted +/- 8 ms to form a fuzzy training set, which was to simulate the possible error in detecting QRS offset point of filtered SAECG. We also found that nonlinear transformation through the hidden layer of developed ANN model could increase Euclidean distance between PEP and NEP patterns.

Algorithms↗

[Mortality of gastric cancer in Catalonia, Spain: geographical distribution and time trends from 1986 to 2000].

BACKGROUND AND OBJECTIVE: The aims of this study are to describe the time trends and the changes in the spatial distribution of stomach cancer mortality by gender, in Catalonia, Spain, in the period 1986-2000. MATERIAL AND METHOD: The mortality data comes from the Mortality Register for Catalonia at the Health Department and the population data from the Institute of Statistics for Catalonia. To analyze time trends, a Poisson regression model was adjusted for each gender. To analyze the geographical distribution, a Bayesian hierarchical model was used. RESULTS: During the period 1986-2000 the number of deaths from stomach cancer was 8,627 for males and 5,831 for females. During this period the estimated decrease in mortality was 3.13% for males and 3.91% for females. The spatial analysis showed the lowest mortality risk areas along the coast while the mortality risk increased toward the zones in the interior. This geographical pattern is very similar for both sexes but in the lasts years of the period it has been fading. CONCLUSIONS: The time trends and the geographical pattern of stomach cancer mortality in Catalonia is similar for both sexes and it is consistent with the trends observed in other developed countries. This suggests a relationship with improved food habits and a better accessibility to health care in the areas of higher risk.

Catchment Area, Health↗

Prognostic meta-signature of breast cancer developed by two-stage mixture modeling of microarray data.

BACKGROUND: An increasing number of studies have profiled tumor specimens using distinct microarray platforms and analysis techniques. With the accumulating amount of microarray data, one of the most intriguing yet challenging tasks is to develop robust statistical models to integrate the findings. RESULTS: By applying a two-stage Bayesian mixture modeling strategy, we were able to assimilate and analyze four independent microarray studies to derive an inter-study validated "meta-signature" associated with breast cancer prognosis. Combining multiple studies (n = 305 samples) on a common probability scale, we developed a 90-gene meta-signature, which strongly associated with survival in breast cancer patients. Given the set of independent studies using different microarray platforms which included spotted cDNAs, Affymetrix GeneChip, and inkjet oligonucleotides, the individually identified classifiers yielded gene sets predictive of survival in each study cohort. The study-specific gene signatures, however, had minimal overlap with each other, and performed poorly in pairwise cross-validation. The meta-signature, on the other hand, accommodated such heterogeneity and achieved comparable or better prognostic performance when compared with the individual signatures. Further by comparing to a global standardization method, the mixture model based data transformation demonstrated superior properties for data integration and provided solid basis for building classifiers at the second stage. Functional annotation revealed that genes involved in cell cycle and signal transduction activities were over-represented in the meta-signature. CONCLUSION: The mixture modeling approach unifies disparate gene expression data on a common probability scale allowing for robust, inter-study validated prognostic signatures to be obtained. With the emerging utility of microarrays for cancer prognosis, it will be important to establish paradigms to meta-analyze disparate gene expression data for prognostic signatures of potential clinical use.

Bayes Theorem↗

Model selection for mixtures of mutagenetic trees.

The evolution of drug resistance in HIV is characterized by the accumulation of resistance-associated mutations in the HIV genome. Mutagenetic trees, a family of restricted Bayesian tree models, have been applied to infer the order and rate of occurrence of these mutations. Understanding and predicting this evolutionary process is an important prerequisite for the rational design of antiretroviral therapies. In practice, mixtures models of K mutagenetic trees provide more flexibility and are often more appropriate for modelling observed mutational patterns. Here, we investigate the model selection problem for K-mutagenetic trees mixture models. We evaluate several classical model selection criteria including cross-validation, the Bayesian Information Criterion (BIC), and the Akaike Information Criterion. We also use the empirical Bayes method by constructing a prior probability distribution for the parameters of a mutagenetic trees mixture model and deriving the posterior probability of the model. In addition to the model dimension, we consider the redundancy of a mixture model, which is measured by comparing the topologies of trees within a mixture model. Based on the redundancy, we propose a new model selection criterion, which is a modification of the BIC. Experimental results on simulated and on real HIV data show that the classical criteria tend to select models with far too many tree components. Only cross-validation and the modified BIC recover the correct number of trees and the tree topologies most of the time. At the same optimal performance, the runtime of the new BIC modification is about one order of magnitude lower. Thus, this model selection criterion can also be used for large data sets for which cross-validation becomes computationally infeasible.

Bayes Theorem↗

Modelling the effects of biological intervention in a dynamical gene network.

Cellular response to environmental and internal signals can be modeled by dynamical gene regulatory networks (GRN). In the literature, three main classes of gene network models can be distinguished: (1) non-quantitative (or data-based) models which do not describe the probability distribution of gene expressions; (2) quantitative models which fully describe the probability distribution of all genes co-expression; and (3) mechanistic models which allow for a causal interpretation of gene interactions. We propose two rigorous frameworks to model gene alteration in a dynamical GRN, depending on whether the network model is quantitative or mechanistic. We explain how these models can be used for design of experiment, or, if additional alteration data are available, for validation purposes or to improve the parameter estimation of the original model. We apply these methods to the Gaussian graphical model, which is quantitative but non-mechanistic, and to mechanistic models of Bayesian networks and penalized linear regression.

Gene Regulatory Networks↗

Bayesian population analysis of a harmonized physiologically based pharmacokinetic model of trichloroethylene and its metabolites.

Bayesian population analysis of a harmonized physiologically based pharmacokinetic (PBPK) model for trichloroethylene (TCE) and its metabolites was performed. In the Bayesian framework, prior information about the PBPK model parameters is updated using experimental kinetic data to obtain posterior parameter estimates. Experimental kinetic data measured in mice, rats, and humans were available for this analysis, and the resulting posterior model predictions were in better agreement with the kinetic data than prior model predictions. Uncertainty in the prediction of the kinetics of TCE, trichloroacetic acid (TCA), and trichloroethanol (TCOH) was reduced, while the kinetics of other key metabolites dichloroacetic acid (DCA), chloral hydrate (CHL), and dichlorovinyl mercaptan (DCVSH) remain relatively uncertain due to sparse kinetic data for use in this analysis. To help focus future research to further reduce uncertainty in model predictions, a sensitivity analysis was conducted to help identify the parameters that have the greatest impact on various internal dose metric predictions. For application to a risk assessment for TCE, the model provides accurate estimates of TCE, TCA, and TCOH kinetics. This analysis provides an important step toward estimating uncertainty of dose-response relationships in noncancer and cancer risk assessment, improving the extrapolation of toxic TCE doses from experimental animals to humans.

Animals↗

Time-dependent pharmacokinetics of cyclosporine (Neoral) in de novo renal transplant patients.

PURPOSE: A model for the large scale temporal trend in the oral bioavailability of microemulsion cyclosporine (Neoral) (CsA) is established, with dependence on post-(renal) transplantation day (PTD). METHODS: Twenty de novo adult renal transplant recipients were monitored for CsA administered orally q12 h. A model development group (11 patients, 315 blood concentration samples) was screened at 2 h (C(2); n = 92), 3 h (C(3); n = 56) and at predose troughs (C(min); n = 167) over periods of up to 75 days. The final model was tested in nine patients with C(min) (n = 580) monitored across 4-5 years. The doses varied between 100 and 538 mg with an apparent hyperbolic trend in C(2)/dose vs. PTD. A nonlinear mixed effects modelling (NONMEM) approach was used to obtain population and individual patient one-compartment pharmacokinetic (PK) parameters for oral CsA, which carry implicit the bioavailability (F). RESULTS: In the final PK model (PK-f) the F was modelled via a simple function for the temporal (days) trend of the bioavailability after transplantation as, F(f) = 1-alpha * exp(-lambda * PTD) resulting in a 28% reduction in the unexplained intra-individual variability. The population PK-f parameters were, for apparent clearance [mean, 95% confidence interval (interindividual CV%)] Cl/F(f) = 17.0 (13.8-20.2) L/h (27%), apparent central compartment volume of distribution, V/F(f) = 134 L (108-160) (28%), and lambda = 0.037/day (0.005-0.069) (120%). The absorption rate k(a) and the parameter alpha were approximated iteratively as 4/h and 0.62 respectively. The PK-f was structurally superior to the base model in explaining part of the within subject (occasion) variability and predicting the exposure surrogates C(2) and C(3). Also, the PK-f was better than the base model with Bayesian fitting of individual profiles in that group. CONCLUSION: The PTD-dependent relative bioavailability model provides a rational means of steering dose titration of CsA in de novo renal transplantation patients by removing the large scale PK adjustment signal, either through nomograms or as a Bayesian prior.

Adult↗

Modelling techniques and their application for monitoring in high dependency environments--learning models.

This paper reviews the use of learning models including Bayesian classifiers and artificial neural networks in monitoring and interpreting biosignals. Generally learning models applied for analysis of biosignals are "black-box' types trained on the basis of measured signals. It is illustrated that the training and application of learning models more or less follow the same sequences. The main focus is the interpretation of electrical signals from the brain (electroencephalogram (EEG) and evoked potentials (EP)). Current analysis of these signals often reveals sudden changes in the EEG or evoked potentials to be the earliest discernible signs of inadequate perfusion of the brain. They may reflect problems such as systemic arterial oxygen desaturation or hypotension arising from other body system failures during critical illness. It is suggested that these brain signals should be recorded in the critical care unit, and that they should form part of the annotated database of biosignals established during the IMPROVE project. This would allow for the development of new methods for on-line warning of impending damage to the central nervous system, such that corrective actions could be taken before permanent damage occurred.

Auscultation↗

Dynamic Bayesian network and nonparametric regression for nonlinear modeling of gene networks from time series gene expression data.

We propose a dynamic Bayesian network and nonparametric regression model for constructing a gene network from time series microarray gene expression data. The proposed method can overcome a shortcoming of the Bayesian network model in the sense of the construction of cyclic regulations. The proposed method can analyze the microarray data as a continuous data and can capture even nonlinear relations among genes. It can be expected that this model will give a deeper insight into complicated biological systems. We also derive a new criterion for evaluating an estimated network from Bayes approach. We conduct Monte Carlo experiments to examine the effectiveness of the proposed method. We also demonstrate the proposed method through the analysis of the Saccharomyces cerevisiae gene expression data.

Bayes Theorem↗

Bayesian analysis of structural equation models with dichotomous variables.

Structural equation modelling has been used extensively in the behavioural and social sciences for studying interrelationships among manifest and latent variables. Recently, its uses have been well recognized in medical research. This paper introduces a Bayesian approach to analysing general structural equation models with dichotomous variables. In the posterior analysis, the observed dichotomous data are augmented with the hypothetical missing values, which involve the latent variables in the model and the unobserved continuous measurements underlying the dichotomous data. An algorithm based on the Gibbs sampler is developed for drawing the parameters values and the hypothetical missing values from the joint posterior distributions. Useful statistics, such as the Bayesian estimates and their standard error estimates, and the highest posterior density intervals, can be obtained from the simulated observations. A posterior predictive p-value is used to test the goodness-of-fit of the posited model. The methodology is applied to a study of hypertensive patient non-adherence to medication.

Algorithms↗

Rating exposure control using Bayesian decision analysis.

A model is presented for applying Bayesian statistical techniques to the problem of determining, from the usual limited number of exposure measurements, whether the exposure profile for a similar exposure group can be considered a Category 0, 1, 2, 3, or 4 exposure. The categories were adapted from the AIHA exposure category scheme and refer to (0) negligible or trivial exposure (i.e., the true X 0.95 < or =1%OEL), (1) highly controlled (i.e., X 0.95 < or =10%OEL), (2) well controlled (i.e., X 0.95 < or =50%OEL), (3) controlled (i.e., X 0.95 < or =100%OEL), or (4) poorly controlled (i.e., X0.95 > or =1%OEL) exposures. Unlike conventional statistical methods applied to exposure data, Bayesian statistical techniques can be adapted to explicitly take into account professional judgment or other sources of information. The analysis output consists of a distribution (i.e., set) of decision probabilities: e.g., 1%, 80%, 12%, 5%, and 2% probability that the exposure profile is a Category 0, 1, 2, 3, or 4 exposure. By inspection of these decision probabilities, rather than the often difficult to interpret point estimates (e.g., the sample 95th percentile exposure) and confidence intervals, a risk manager can be better positioned to arrive at an effective (i.e., correct) and efficient decision. Bayesian decision methods are based on the concepts of prior, likelihood, and posterior distributions of decision probabilities. The prior decision distribution represents what an industrial hygienist knows about this type of operation, using professional judgment; company, industry, or trade organization experience; historical or surrogate exposure data; or exposure modeling predictions. The likelihood decision distribution represents the decision probabilities based on an analysis of only the current data. The posterior decision distribution is derived by mathematically combining the functions underlying the prior and likelihood decision distributions, and represents the final decision probabilities. Advantages of Bayesian decision analysis include: (a) decision probabilities are easier to understand by risk managers and employees; (b) prior data, professional judgment, or modeling information can be objectively incorporated into the decision-making process; (c) decisions can be made with greater certainty; (d) the decision analysis can be constrained to a more realistic "parameter space" (i.e., the range of plausible values for the true geometric mean and geometric standard deviation); and (e) fewer measurements are necessary whenever the prior distribution is well defined and the process is fairly stable. Furthermore, Bayesian decision analysis provides an obvious feedback mechanism that can be used by an industrial hygienist to improve professional judgment. For example, if the likelihood decision distribution is inconsistent with the prior decision distribution then it is likely that either a significant process change has occurred or the industrial hygienist's initial judgment was incorrect. In either case, the industrial hygienist should readjust his judgment regarding this operation.

Bayes Theorem↗

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem↗

Bayesian analysis of structural equation models with multinomial variables and an application to type 2 diabetic nephropathy.

There is now increasing evidence proving that many complex diseases can be significantly influenced by correlated phenotype and genotype variables, as well as their interactions. Effective and rigorous assessment of such influence is difficult, because the number of phenotype and genotype variables of interest may not be small, and a genotype variable is an unordered categorical variable that follows a multinomial distribution. To address the problem, we establish a novel nonlinear structural equation model for analysing mixed continuous and multinomial data that can be missing at random. A confirmatory factor analysis model with Kronecker product is proposed for grouping the manifest continuous and multinomial variables into latent variables according to their functions; and a nonlinear structural equation is formulated to assess the linear and interaction effects of the independent latent variables to the dependent latent variables. Bayesian methods for estimation and model comparison are developed through Markov chain Monte Carlo techniques and path sampling. The newly developed methodologies are applied to a case-control cohort of type 2 diabetic patients with nephropathy.

Bayes Theorem↗

Comparison of linear and tapered intravenous infusion of methotrexate in oncochemotherapy. A theoretical approach.

In oncochemotherapy with methotrexate (MTX) a peripheral concentration greater than 0.45 mg/l and a plasma concentration less than 45 mg/l must be maintained for 20 h. The time periods required to reach and maintain steady-state concentrations after tapered and linear intravenous infusion were compared. Pharmacokinetic analyses according to a two-compartment model were used to calculate dosage regimens and concentration profiles by means of the Bayesian General Modelling Program (BM) and NONLIN. When the dosage regimen is based on a steady-state concentration in the peripheral compartment (which is the target compartment for MTX) tapered infusion reaches this concentration 40% faster and maintains it 12.5% longer, but no difference is found if the dosage regimen is based on a steady-state concentration in the central compartment. In theory the two-step 24-hour tapered infusion can be replaced by a bolus injection plus linear infusion in the ratio 1:2 of the total dose. These dosage regimens are to be preferred over linear infusion.

Humans↗

Flexible empirical Bayes models for differential gene expression.

MOTIVATION: Inference about differential expression is a typical objective when analyzing gene expression data. Recently, Bayesian hierarchical models have become increasingly popular for this type of problem. The two most common hierarchical models are the hierarchical Gamma-Gamma (GG) and Lognormal-Normal (LNN) models. However, to facilitate inference, some unrealistic assumptions have been made. One such assumption is that of a common coefficient of variation across genes, which can adversely affect the resulting inference. RESULTS: In this paper, we extend both the GG and LNN modeling frameworks to allow for gene-specific variances and propose EM based algorithms for parameter estimation. The proposed methodology is evaluated on three experimental datasets: one cDNA microarray experiment and two Affymetrix spike-in experiments. The two extended models significantly reduce the false positive rate while keeping a high sensitivity when compared to the originals. Finally, using a simulation study we show that the new frameworks are also more robust to model misspecification. AVAILABILITY: The R code for implementing the proposed methodology can be downloaded at http://www.stat.ubc.ca/~c.lo/FEBarrays. SUPPLEMENTARY INFORMATION: The supplementary material is available at http://www.stat.ubc.ca/~c.lo/FEBarrays/supp.pdf.

Algorithms↗