Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Bayesian neural network approaches to ovarian cancer identification from high-resolution mass spectrometry data.

MOTIVATION: The classification of high-dimensional data is always a challenge to statistical machine learning. We propose a novel method named shallow feature selection that assigns each feature a probability of being selected based on the structure of training data itself. Independent of particular classifiers, the high dimension of biodata can be fleetly reduced to an applicable case for consequential processing. Moreover, to improve both efficiency and performance of classification, these prior probabilities are further used to specify the distributions of top-level hyperparameters in hierarchical models of Bayesian neural network (BNN), as well as the parameters in Gaussian process models. RESULTS: Three BNN approaches were derived and then applied to identify ovarian cancer from NCI's high-resolution mass spectrometry data, which yielded an excellent performance in 1000 independent k-fold cross validations (k = 2,...,10). For instance, indices of average sensitivity and specificity of 98.56 and 98.42%, respectively, were achieved in the 2-fold cross validations. Furthermore, only one control and one cancer were misclassified in the leave-one-out cross validation. Some other popular classifiers were also tested for comparison. AVAILABILITY: The programs implemented in MatLab, R and Neal's fbm.2004-11-10.

Bayes Theorem↗

A Bayesian compound stochastic process for modeling nonstationary and nonhomogeneous sequence evolution.

Variations of nucleotidic composition affect phylogenetic inference conducted under stationary models of evolution. In particular, they may cause unrelated taxa sharing similar base composition to be grouped together in the resulting phylogeny. To address this problem, we developed a nonstationary and nonhomogeneous model accounting for compositional biases. Unlike previous nonstationary models, which are branchwise, that is, assume that base composition only changes at the nodes of the tree, in our model, the process of compositional drift is totally uncoupled from the speciation events. In addition, the total number of events of compositional drift distributed across the tree is directly inferred from the data. We implemented the method in a Bayesian framework, relying on Markov Chain Monte Carlo algorithms, and applied it to several nucleotidic data sets. In most cases, the stationarity assumption was rejected in favor of our nonstationary model. In addition, we show that our method is able to resolve a well-known artifact. By Bayes factor evaluation, we compared our model with 2 previously developed nonstationary models. We show that the coupling between speciations and compositional shifts inherent to branchwise models may lead to an overparameterization, resulting in a lesser fit. In some cases, this leads to incorrect conclusions, concerning the nature of the compositional biases. In contrast, our compound model more flexibly adapts its effective number of parameters to the data sets under investigation. Altogether, our results show that accounting for nonstationary sequence evolution may require more elaborate and more flexible models than those currently used.

Animals↗

A context noise model of episodic word recognition.

Item noise models of recognition assert that interference at retrieval is generated by the words from the study list. Context noise models of recognition assert that interference at retrieval is generated by the contexts in which the test word has appeared. The authors introduce the bind cue decide model of episodic memory, a Bayesian context noise model, and demonstrate how it can account for data from the item noise and dual-processing approaches to recognition memory. From the item noise perspective, list strength and list length effects, the mirror effect for word frequency and concreteness, and the effects of the similarity of other words in a list are considered. From the dual-processing perspective, process dissociation data on the effects of length, temporal separation of lists, strength, and diagnosticity of context are examined. The authors conclude that the context noise approach to recognition is a viable alternative to existing approaches.

Attention↗

H-CORE: enabling genome-scale Bayesian analysis of biological systems without prior knowledge.

The Bayesian network is a popular tool for describing relationships between data entities by representing probabilistic (in)dependencies with a directed acyclic graph (DAG) structure. Relationships have been inferred between biological entities using the Bayesian network model with high-throughput data from biological systems in diverse fields. However, the scalability of those approaches is seriously restricted because of the huge search space for finding an optimal DAG structure in the process of Bayesian network learning. For this reason, most previous approaches limit the number of target entities or use additional knowledge to restrict the search space. In this paper, we use the hierarchical clustering and order restriction (H-CORE) method for the learning of large Bayesian networks by clustering entities and restricting edge directions between those clusters, with the aim of overcoming the scalability problem and thus making it possible to perform genome-scale Bayesian network analysis without additional biological knowledge. We use simulations to show that H-CORE is much faster than the widely used sparse candidate method, whilst being of comparable quality. We have also applied H-CORE to retrieving gene-to-gene relationships in a biological system (The 'Rosetta compendium'). By evaluating learned information through literature mining, we demonstrate that H-CORE enables the genome-scale Bayesian analysis of biological systems without any prior knowledge.

Algorithms↗

Broadening the Tests of Learning Models.

For many years psychological studies of the learning process have used a simulated medical diagnosis task in which symptom configurations are probabilistically related to diseases. Participants are given a set of symptoms and asked to indicate which disease is present, and feedback is given on each trial. We enrich this standard laboratory task in four different ways. First, the symptoms have four possible values (low, medium low, medium high, and high) rather than just two. Second, symptom configurations are generated from an expanded factorial design rather than a simple factorial design. Third, subjects are asked to make a continuous judgment indicating their confidence in the diagnosis, rather than simply a binary judgment. Fourth, cumulated performance scores, payoffs, and the availability of a historical summary of the outcomes are varied in order to assess how these treatments modulate performance. These enrichments provide a broader data set and more challenging tests of the models. Using 123 subjects each in 480 trials, we compare five existing learning models plus several variants, including the well-known Bayesian, fuzzy logic, connectionist, exemplar, and ALCOVE models. We find that the subjects do learn to distinguish the symptom configurations, that subjects are quite heterogeneous in their response to the task, and that only a small part of the variation across subjects arises from the differences in treatments. The most striking finding is that the model that best predicts subjects' behavior is a simple Bayesian model with a single fitted parameter for prior precision to capture individual differences. We use rolling regression techniques to elucidate the behavior of this model over time and find some evidence of over-response to current stimuli. Copyright 1998 Academic Press.

Journal Article↗

Pathways of urothelial cancer progression suggested by Bayesian network analysis of allelotyping data.

Urothelial cancers of the bladder (UC) comprise biologically heterogeneous group of tumors and display complex genetic alterations. Several genetic changes have been analyzed in detail and some of them are associated with the development and progression of UCs. Only a few studies, however, are focused on identifying the order in which the aberrations may appear during UC tumorigenesis. We have analyzed 123 papillary UCs of the bladder by microsatellites for each of the chromosomal regions that have been suggested to be specifically involved in this type of tumor. We used Bayesian network modeling that enables to uncover multivariate probabilistic dependencies between variables. This methodology applied to LOH data allowed us to discover patterns of losses in UCs. Exploiting the mechanism of probabilistic reasoning in Bayesian networks we suggest primary and secondary events in tumor pathogenesis and reconstruct the possible flow of progression of allelic changes. Losses of chromosome 9p and 9q were found to be the primary events. Losses of 8p and 17p are important events leading to progression of tumor cell clones. The loss of 17p occurs when both abnormalities of chromosome 9 and 8p are already present. There are chromosomal losses related to 8p (1q, 18q, 10q) and some losses like 5q/5p were associated with 17p, leading to the hypothesis of different genetic pathways of UC progression. The abnormalities of chromosome regions 13q, 16q, 6q, 14q, 3p are suggested to be late events being accumulated during the progression of cancer. Although some genetic changes were associated only with the 8p pathway, most secondary genetic changes appear in both pathways. Supplementary material for this article can be found on the International Journal of Cancer website at http://www.interscience.wiley.com/jpages/0020-7136/suppmat/index.html.

Alleles↗

A stochastic method for Bayesian estimation of hidden Markov random field models with application to a color model.

We propose a new stochastic algorithm for computing useful Bayesian estimators of hidden Markov random field (HMRF) models that we call exploration/selection/estimation (ESE) procedure. The algorithm is based on an optimization algorithm of O. François, called the exploration/selection (E/S) algorithm. The novelty consists of using the a posteriori distribution of the HMRF, as exploration distribution in the E/S algorithm. The ESE procedure computes the estimation of the likelihood parameters and the optimal number of region classes, according to global constraints, as well as the segmentation of the image. In our formulation, the total number of region classes is fixed, but classes are allowed or disallowed dynamically. This framework replaces the mechanism of the split-and-merge of regions that can be used in the context of image segmentation. The procedure is applied to the estimation of a HMRF color model for images, whose likelihood is based on multivariate distributions, with each component following a Beta distribution. Meanwhile, a method for computing the maximum likelihood estimators of Beta distributions is presented. Experimental results performed on 100 natural images are reported. We also include a proof of convergence of the E/S algorithm in the case of nonsymmetric exploration graphs.

Algorithms↗

Applying a moving total mortality count to the cities in the NMMAPS database to estimate the mortality effects of particulate matter air pollution.

OBJECTIVES: To apply a new method for estimating the association between daily ambient particulate matter air pollution (PM) and daily mortality to data from over 100 United States cities contained in the National Morbidity, Mortality, and Air Pollution Study (NMMAPS) database and to see whether the results from the 90 cities NMMAPS analysis are robust to this different modelling approach. This new method has recently been shown to provide improved estimates for the association between PM and daily mortality when every-day PM data are unavailable. It avoids the need for selecting a lag of PM at which the mortality effects of PM are to be investigated. METHODS: With the aid of analytical methods and databases developed for NMMAPS, Poisson log linear models controlling for long term trends and weather effects were used to estimate the association between PM and mortality for cities in the NMMAPS database using the new method. A two stage Bayesian hierarchical model was then used to combine city specific estimates to form a national average PM mortality effect estimate. RESULTS: A 10 microg/m3 increase in PM was associated with a 0.12% increment in total mortality and a 0.17% increment in cardiovascular and respiratory mortality. These results are consistent with those found in the NMMAPS analysis. CONCLUSIONS: There is a statistically significant association between short term changes in PM and mortality on average for the cities contained in the NMMAPS database. These findings are further evidence that this widespread pollutant adversely affects public health.

Adult↗

An empirical Bayes approach to inferring large-scale gene association networks.

MOTIVATION: Genetic networks are often described statistically using graphical models (e.g. Bayesian networks). However, inferring the network structure offers a serious challenge in microarray analysis where the sample size is small compared to the number of considered genes. This renders many standard algorithms for graphical models inapplicable, and inferring genetic networks an 'ill-posed' inverse problem. METHODS: We introduce a novel framework for small-sample inference of graphical models from gene expression data. Specifically, we focus on the so-called graphical Gaussian models (GGMs) that are now frequently used to describe gene association networks and to detect conditionally dependent genes. Our new approach is based on (1) improved (regularized) small-sample point estimates of partial correlation, (2) an exact test of edge inclusion with adaptive estimation of the degree of freedom and (3) a heuristic network search based on false discovery rate multiple testing. Steps (2) and (3) correspond to an empirical Bayes estimate of the network topology. RESULTS: Using computer simulations, we investigate the sensitivity (power) and specificity (true negative rate) of the proposed framework to estimate GGMs from microarray data. This shows that it is possible to recover the true network topology with high accuracy even for small-sample datasets. Subsequently, we analyze gene expression data from a breast cancer tumor study and illustrate our approach by inferring a corresponding large-scale gene association network for 3883 genes.

Algorithms↗

A model-based method for identifying species hybrids using multilocus genetic data.

We present a statistical method for identifying species hybrids using data on multiple, unlinked markers. The method does not require that allele frequencies be known in the parental species nor that separate, pure samples of the parental species be available. The method is suitable for both markers with fixed allelic differences between the species and markers without fixed differences. The probability model used is one in which parentals and various classes of hybrids (F(1)'s, F(2)'s, and various backcrosses) form a mixture from which the sample is drawn. Using the framework of Bayesian model-based clustering allows us to compute, by Markov chain Monte Carlo, the posterior probability that each individual belongs to each of the distinct hybrid classes. We demonstrate the method on allozyme data from two species of hybridizing trout, as well as on two simulated data sets.

Animals↗

Bayesian 2-D deconvolution: effect of using spatially invariant ultrasound point spread functions.

Observed ultrasound images are degraded representations of the-true tissue reflectance. The specular reflections at boundaries between regions of different tissue types are blurred, and the diffuse scattering within homogenous regions causes speckle because of the oscillating nature of the transmitted pulse. To reduce both blur and speckle, we have developed algorithms for the restoration of simulated and real ultrasound images based on Markov random field models and Bayesian statistical methods. The algorithm is summarized here, although a more detailed description can be found in our companion paper [1]. Because the point spread function (psf) is unknown, we investigate the effects of using incorrect frequencies and sizes for the model psf during the restoration process. First, we degrade the images either with a known simulated psf or a measured psf. Then, we use different psf shapes during restoration to study the robustness of the method. We found that small variations in the parameters characterizing the psf, less than +/- 25% change in frequency, width, or length, still yielded satisfactory results. When altering the psf more than this, the restorations were not acceptable. The restorations were particularly sensitive to large increases in the restoring psf frequency. Thus, 2-D Bayesian restoration using a fixed psf may yield acceptable results as long as the true variant psfs have not varied too much during imaging.

Bayes Theorem↗

Prediction of the international normalized ratio and maintenance dose during the initiation of warfarin therapy.

AIMS: A pharmacokinetic/pharmacodynamic model, with Bayesian parameter estimation, was used to retrospectively predict the daily International Normalized Ratios (INRs) and the maintenance doses during the initiation of warfarin therapy in 74 inpatients. METHODS: INRs and maintenance doses predicted by the model were compared with the actual INRs and the eventual maintenance dose. Cases with drugs or medical conditions interacting with warfarin or receiving concurrent heparin therapy were not excluded. As the study was retrospective, model predictions of the maintenance dose were not those that were administered. Mean prediction error (MPE) and percentage absolute prediction errors (PAPE) were used to assess the model predictions. RESULTS: INR MPE ranged from -0.07 to 0.06 and median PAPE from 10% to 20%. Dose MPE ranged from -0.7 to 0.17 mg and median PAPE from 16.7% to 37.5%. Accurate and precise dose predictions were obtained after 3 or more INR feedback's. CONCLUSIONS: This study shows that the model can accurately predict daily INRs and the maintenance dose in this sample of cases. The model can be incorporated into computer decision-support systems for warfarin therapy and may lead to improvement in the initiation of warfarin therapy.

Adult↗

A hierarchical Naïve Bayes Model for handling sample heterogeneity in classification problems: an application to tissue microarrays.

BACKGROUND: Uncertainty often affects molecular biology experiments and data for different reasons. Heterogeneity of gene or protein expression within the same tumor tissue is an example of biological uncertainty which should be taken into account when molecular markers are used in decision making. Tissue Microarray (TMA) experiments allow for large scale profiling of tissue biopsies, investigating protein patterns characterizing specific disease states. TMA studies deal with multiple sampling of the same patient, and therefore with multiple measurements of same protein target, to account for possible biological heterogeneity. The aim of this paper is to provide and validate a classification model taking into consideration the uncertainty associated with measuring replicate samples. RESULTS: We propose an extension of the well-known Naïve Bayes classifier, which accounts for biological heterogeneity in a probabilistic framework, relying on Bayesian hierarchical models. The model, which can be efficiently learned from the training dataset, exploits a closed-form of classification equation, thus providing no additional computational cost with respect to the standard Naïve Bayes classifier. We validated the approach on several simulated datasets comparing its performances with the Naïve Bayes classifier. Moreover, we demonstrated that explicitly dealing with heterogeneity can improve classification accuracy on a TMA prostate cancer dataset. CONCLUSION: The proposed Hierarchical Naïve Bayes classifier can be conveniently applied in problems where within sample heterogeneity must be taken into account, such as TMA experiments and biological contexts where several measurements (replicates) are available for the same biological sample. The performance of the new approach is better than the standard Naïve Bayes model, in particular when the within sample heterogeneity is different in the different classes.

Algorithms↗

Extended-interval dosing of tobramycin in neonates: implications for therapeutic drug monitoring.

OBJECTIVE: Our objective was to individualize tobramycin dosing regimens in neonates of various gestational ages with use of early therapeutic drug monitoring. METHODS: This study was performed in neonatal patients with suspected septicemia in the first week of life. All patients received tobramycin, 4 mg/kg per dose, as a 30-minute intravenous infusion, with a gestational age-related initial interval of 48 hours (<32 weeks), 36 hours (32-36 weeks), and 24 hours (> or =37 weeks). The target serum peak and trough serum concentrations were 5 to 10 mg/L and 0.5 mg/L, respectively. Serum trough samples and 1- and 6-hour samples were taken after the first dose. Tobramycin concentrations were used to obtain gestational age-dependent population models with nonparametric expectation maximization software. To investigate the effect of timing of sampling in a second group of patients, serum trough samples and 3- and 8-hour samples were taken after the first dose of tobramycin was administered. Serum trough concentrations were predicted by use of linear pharmacokinetics in both groups and by use of the population models with bayesian feedback of 1 or 2 serum concentrations in the second group. These predicted concentrations were compared with actual serum trough concentrations. The predictive performance of the 1- to 6-hour and 3- to 8-hour models and the population models were compared with a gestational age-related model without therapeutic drug monitoring. RESULTS: A total of 247 patients were analyzed: 206 with 1- to 6-hour serum samples and 41 with 3- to 8-hour serum samples. Peak serum concentrations were above 5 mg/L in 90.8% of cases, and trough serum concentrations were above 1 mg/L in 25.5% of cases. The 3- to 8-hour linear model had a bias of -0.31 mg/L and a precision of 0.48 mg/L, and it performed significantly better than the 1- to 6-hour model. The best nonparametric expectation maximization model had a bias of -0.11 mg/L and a precision of 0.45 mg/L. None of the models yielded a significant improvement of predictive performance over the model without therapeutic drug monitoring. CONCLUSIONS: Routine early therapeutic drug monitoring does not improve the model-based prediction of initial tobramycin dosing intervals in neonates in the first week of life.

Anti-Bacterial Agents↗

Semiparametric proportional odds models for spatially correlated survival data.

The last decade has witnessed major developments in Geographical Information Systems (GIS) technology resulting in the need for statisticians to develop models that account for spatial clustering and variation. In public health settings, epidemiologists and health-care professionals are interested in discerning spatial patterns in survival data that might exist among the counties. This paper develops a Bayesian hierarchical model for capturing spatial heterogeneity within the framework of proportional odds. This is deemed more appropriate when a substantial percentage of subjects enjoy prolonged survival. We discuss the implementation issues of our models, perform comparisons among competing models and illustrate with data from the SEER (Surveillance Epidemiology and End Results) database of the National Cancer Institute, paying particular attention to the underlying spatial story.

Bayes Theorem↗

Ensembles of Bayesian-regularized genetic neural networks for modeling of acetylcholinesterase inhibition by huprines.

Acetylcholinesterase inhibition was modeled for a set of huprines using ensembles of Bayesian-regularized Genetic Neural Networks. In the Bayesian-regularized Genetic Neural Network approach the Bayesian regularization avoids overfitted regressions and the genetic algorithm allows exploring a wide pool of three-dimensional descriptors. The predictive capacity of our selected model was evaluated by averaging multiple validation sets generated as members of neural network ensembles. When 60 members are assembled, the neural network ensemble provides a reliable measure of training and test set R(2)-values of 0.945 and 0.850 respectively. In other respects, the ability of the nonlinear selected genetic algorithm space for differentiate the data were evidenced when total data set was well distributed in a Kohonen self-organizing map. The analysis of the self-organizing map zones allows establishing the main structural features differentiated by our vectorial space.

Acetylcholinesterase↗

Empirical Bayes versus fully Bayesian analysis of geographical variation in disease risk.

This paper reviews methods for mapping geographical variation in disease incidence and mortality. Recent results in Bayesian hierarchical modelling of relative risk are discussed. Two approaches to relative risk estimation, along with the related computational procedures, are described and compared. The first is an empirical Bayes approach that uses a technique of penalized log-likelihood maximization; the second approach is fully Bayesian, and uses an innovative stochastic simulation technique called the Gibbs sampler. We chose to map geographical variation in breast cancer and Hodgkin's disease mortality as observed in all the health care districts of Sardinia, to illustrate relevant problems, methods and techniques.

Bayes Theorem↗

Bayesian analysis of a dose-response experiment with serial sacrifices.

This paper presents analysis and comments which are believed to be appropriate for certain carcinogenesis studies where sacrifices are performed throughout the experiment. Estimates of the risk probability for each dose level and sacrifice time are found utilizing the sample likelihood as the posterior density. The dose-response relationship is investigated with these estimates as the response. In order to test if the dose is effective and to check the appropriateness of the time-to-incidence model a Bayesian multiple comparisons technique is introduced.

Animals↗