Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Processing radio frequency ultrasound images: a robust method for local spectral features estimation by a spatially constrained parametric approach.

Spectral estimation is a major component in studies aiming at characterizing biological tissues through the analysis of backscattered radio frequency (RF) ultrasonic signals and images. However, conventional spectral estimation techniques yield a well-known trade-off between spatial resolution and variance. The backscattered signals are stochastic by nature, so short-term local analysis results in a high variance of the estimates, which cannot efficiently be reduced through conventional spatial averaging. We address this issue by describing a spectral estimation technique that reduces the variance of the estimates (by smoothing the local estimates in spectrally homogeneous regions) while preserving spectral discontinuities (i.e., the smoothing is not performed across regions with different spectral contents). The proposed approach is set in a Bayesian framework and is based on local autoregressive (AR) estimation, constrained by smoothness priors. These smoothness priors are introduced through a Markov random field in which the associated potential functions are nonquadratic, allowing thereby to preserve discontinuity. The method is validated on simulated RF images and tested on echocardiographic images acquired in vivo. The results are compared to the estimates provided by the conventional Burg technique. These results clearly demonstrate the ability of the proposed approach to improve spectral estimation in terms of variance reduction and discontinuity detection.

Algorithms↗

Identifying protein complexes in high-throughput protein interaction screens using an infinite latent feature model.

We propose a Bayesian approach to identify protein complexes and their constituents from high-throughput protein-protein interaction screens. An infinite latent feature model that allows for multi-complex membership by individual proteins is coupled with a graph diffusion kernel that evaluates the likelihood of two proteins belonging to the same complex. Gibbs sampling is then used to infer a catalog of protein complexes from the interaction screen data. An advantage of this model is that it places no prior constraints on the number of complexes and automatically infers the number of significant complexes from the data. Validation results using affinity purification/mass spectrometry experimental data from yeast RNA-processing complexes indicate that our method is capable of partitioning the data in a biologically meaningful way. A supplementary web site containing larger versions of the figures is available at http://public.kgi.edu/wild/PSBO6/index.html.

Algorithms↗

Bayesian methods for cluster randomized trials with continuous responses.

Bayesian methods for cluster randomized trials extend the random-effects formulation by allowing both the use of external evidence on parameters and straightforward relaxation of the standard normality and constant variance assumptions. Care is required in specifying prior distributions on variance components, and a number of different options are explored with implied prior distributions for other parameters given in closed form. Markov chain Monte Carlo (MCMC) methods permit the fitting of very general models and the introduction of parameter uncertainty into power calculations. We illustrate these ideas using a published example in which general practices were randomized to intervention or control, and show that different choices of supposedly 'non-informative' prior distributions can have substantial influence on conclusions. We also illustrate the use of forward simulation methods in power calculations with uncertainty on multiple inputs. Bayesian methods have the potential to be very useful but guidance is required as to appropriate strategies for robust analysis. Our current experience leads us to recommend a standard 'non-informative' prior distribution for the within-cluster sampling variance, and an independent prior on the intraclass correlation coefficient (ICC). The latter may exploit background evidence or, as a reference analysis, be a uniform ICC or a 'uniform shrinkage' prior.

Bayes Theorem↗

Implementing the Bayesian paradigm: reporting research results over the World-Wide Web.

For decades, statisticians, philosophers, medical investigators and others interested in data analysis have argued that the Bayesian paradigm is the proper approach for reporting the results of scientific analyses for use by clients and readers. To date, the methods have been too complicated for non-statisticians to use. In this paper we argue that the World-Wide Web provides the perfect environment to put the Bayesian paradigm into practice: the likelihood function of the data is parsimoniously represented on the server side, the reader uses the client to represent her prior belief, and a downloaded program (a Java applet) performs the combination. In our approach, a different applet can be used for each likelihood function, prior belief can be assessed graphically, and calculation results can be reported in a variety of ways. We present a prototype implementation, BayesApplet, for two-arm clinical trials with normally-distributed outcomes, a prominent model for clinical trials. The primary implication of this work is that publishing medical research results on the Web can take a form beyond or different from that currently used on paper, and can have a profound impact on the publication and use of research results.

Bayes Theorem↗

Bayesian model accounting for within-class biological variability in Serial Analysis of Gene Expression (SAGE).

BACKGROUND: An important challenge for transcript counting methods such as Serial Analysis of Gene Expression (SAGE), "Digital Northern" or Massively Parallel Signature Sequencing (MPSS), is to carry out statistical analyses that account for the within-class variability, i.e., variability due to the intrinsic biological differences among sampled individuals of the same class, and not only variability due to technical sampling error. RESULTS: We introduce a Bayesian model that accounts for the within-class variability by means of mixture distribution. We show that the previously available approaches of aggregation in pools ("pseudo-libraries") and the Beta-Binomial model, are particular cases of the mixture model. We illustrate our method with a brain tumor vs. normal comparison using SAGE data from public databases. We show examples of tags regarded as differentially expressed with high significance if the within-class variability is ignored, but clearly not so significant if one accounts for it. CONCLUSION: Using available information about biological replicates, one can transform a list of candidate transcripts showing differential expression to a more reliable one. Our method is freely available, under GPL/GNU copyleft, through a user friendly web-based on-line tool or as R language scripts at supplemental web-site.

Astrocytoma↗

Sample size calculations for surveys to substantiate freedom of populations from infectious agents.

We develop a Bayesian approach to sample size computations for surveys designed to provide evidence of freedom from a disease or from an infectious agent. A population is considered "disease-free" when the prevalence or probability of disease is less than some threshold value. Prior distributions are specified for diagnostic test sensitivity and specificity and we test the null hypothesis that the prevalence is below the threshold. Sample size computations are developed using hypergeometric sampling for finite populations and binomial sampling for infinite populations. A normal approximation is also developed. Our procedures are compared with the frequentist methods of Cameron and Baldock (1998a, Preventive Veterinary Medicine34, 1-17.) using an example of foot-and-mouth disease. User-friendly programs for sample size calculation and analysis of survey data are available at http://www.epi.ucdavis.edu/diagnostictests/.

Animals↗

Brain mechanism of reward prediction under predictable and unpredictable environmental dynamics.

In learning goal-directed behaviors, an agent has to consider not only the reward given at each state but also the consequences of dynamic state transitions associated with action selection. To understand brain mechanisms for action learning under predictable and unpredictable environmental dynamics, we measured brain activities by functional magnetic resonance imaging (fMRI) during a Markov decision task with predictable and unpredictable state transitions. Whereas the striatum and orbitofrontal cortex (OFC) were significantly activated both under predictable and unpredictable state transition rules, the dorsolateral prefrontal cortex (DLPFC) was more strongly activated under predictable than under unpredictable state transition rules. We then modelled subjects' choice behaviours using a reinforcement learning model and a Bayesian estimation framework and found that the subjects took larger temporal discount factors under predictable state transition rules. Model-based analysis of fMRI data revealed different engagement of striatum in reward prediction under different state transition dynamics. The ventral striatum was involved in reward prediction under both unpredictable and predictable state transition rules, although the dorsal striatum was dominantly involved in reward prediction under predictable rules. These results suggest different learning systems in the cortico-striatum loops depending on the dynamics of the environment: the OFC-ventral striatum loop is involved in action learning based on the present state, while the DLPFC-dorsal striatum loop is involved in action learning based on predictable future states.

Brain↗

Some extensions and applications of a Bayesian strategy for monitoring multiple outcomes in clinical trials.

We present some practical extensions and applications of a strategy proposed by Thall, Simon and Estey for designing and monitoring single-arm clinical trials with multiple outcomes. We show by application how the strategy may be applied to construct designs for phase IIA activity trials and phase II equivalence trials. We also show how it may be extended to incorporate the use of mixture priors in settings where a Dirichlet distribution does not adequately quantify prior experience, randomized phase II selection trials involving two or more experimental treatments, and trials with group-sequential monitoring for applications involving multiple institutions.

Antineoplastic Agents↗

Retinal vessel segmentation using the 2-D Gabor wavelet and supervised classification.

We present a method for automated segmentation of the vasculature in retinal images. The method produces segmentations by classifying each image pixel as vessel or nonvessel, based on the pixel's feature vector. Feature vectors are composed of the pixel's intensity and two-dimensional Gabor wavelet transform responses taken at multiple scales. The Gabor wavelet is capable of tuning to specific frequencies, thus allowing noise filtering and vessel enhancement in a single step. We use a Bayesian classifier with class-conditional probability density functions (likelihoods) described as Gaussian mixtures, yielding a fast classification, while being able to model complex decision surfaces. The probability distributions are estimated based on a training set of labeled pixels obtained from manual segmentations. The method's performance is evaluated on publicly available DRIVE (Staal et al., 2004) and STARE (Hoover et al., 2000) databases of manually labeled images. On the DRIVE database, it achieves an area under the receiver operating characteristic curve of 0.9614, being slightly superior than that presented by state-of-the-art approaches. We are making our implementation available as open source MATLAB scripts for researchers interested in implementation details, evaluation, or development of methods.

Algorithms↗

Probabilistic annotation of protein sequences based on functional classifications.

BACKGROUND: One of the most evident achievements of bioinformatics is the development of methods that transfer biological knowledge from characterised proteins to uncharacterised sequences. This mode of protein function assignment is mostly based on the detection of sequence similarity and the premise that functional properties are conserved during evolution. Most automatic approaches developed to date rely on the identification of clusters of homologous proteins and the mapping of new proteins onto these clusters, which are expected to share functional characteristics. RESULTS: Here, we inverse the logic of this process, by considering the mapping of sequences directly to a functional classification instead of mapping functions to a sequence clustering. In this mode, the starting point is a database of labelled proteins according to a functional classification scheme, and the subsequent use of sequence similarity allows defining the membership of new proteins to these functional classes. In this framework, we define the Correspondence Indicators as measures of relationship between sequence and function and further formulate two Bayesian approaches to estimate the probability for a sequence of unknown function to belong to a functional class. This approach allows the parametrisation of different sequence search strategies and provides a direct measure of annotation error rates. We validate this approach with a database of enzymes labelled by their corresponding four-digit EC numbers and analyse specific cases. CONCLUSION: The performance of this method is significantly higher than the simple strategy consisting in transferring the annotation from the highest scoring BLAST match and is expected to find applications in automated functional annotation pipelines.

Algorithms↗

Detection of multiple QTL with epistatic effects under a mixed inheritance model in an outbred population.

A quantitative trait depends on multiple quantitative trait loci (QTL) and on the interaction between two or more QTL, named epistasis. Several methods to detect multiple QTL in various types of design have been proposed, but most of these are based on the assumption that each QTL works independently and epistasis has not been explored sufficiently. The objective of the study was to propose an integrated method to detect multiple QTL with epistases using Bayesian inference via a Markov chain Monte Carlo (MCMC) algorithm. Since the mixed inheritance model is assumed and the deterministic algorithm to calculate the probabilities of QTL genotypes is incorporated in the method, this can be applied to an outbred population such as livestock. Additionally, we treated a pair of QTL as one variable in the Reversible jump Markov chain Monte Carlo (RJMCMC) algorithm so that two QTL were able to be simultaneously added into or deleted from a model. As a result, both of the QTL can be detected, not only in cases where either of the two QTL has main effects and they have epistatic effects between each other, but also in cases where neither of the two QTL has main effects but they have epistatic effects. The method will help ascertain the complicated structure of quantitative traits.

Algorithms↗

Improved pairwise alignments of proteins in the Twilight Zone using local structure predictions.

MOTIVATION: In recent years, advances have been made in the ability of computational methods to discriminate between homologous and non-homologous proteins in the 'twilight zone' of sequence similarity, where the percent sequence identity is a poor indicator of homology. To make these predictions more valuable to the protein modeler, they must be accompanied by accurate alignments. Pairwise sequence alignments are inferences of orthologous relationships between sequence positions. Evolutionary distance is traditionally modeled using global amino acid substitution matrices. But real differences in the likelihood of substitutions may exist for different structural contexts within proteins, since structural context contributes to the selective pressure. RESULTS: HMMSUM (HMMSTR-based substitution matrices) is a new model for structural context-based amino acid substitution probabilities consisting of a set of 281 matrices, each for a different sequence-structure context. HMMSUM does not require the structure of the protein to be known. Instead, predictions of local structure are made using HMMSTR, a hidden Markov model for local structure. Alignments using the HMMSUM matrices compare favorably to alignments carried out using the BLOSUM matrices or structure-based substitution matrices SDM and HSDM when validated against remote homolog alignments from BAliBASE. HMMSUM has been implemented using local Dynamic Programming and with the Bayesian Adaptive alignment method.

Algorithms↗

Automated classification of clustered microcalcifications into malignant and benign types.

The objectives in this study were to design and test a fully automated method for classification of microcalcification clusters into malignant and benign types, and to compare the method's performance with that of radiologists. A novel aspect of the approach is that the relative location and orientation of clusters inside the breast was taken into account for feature calculation. Furthermore, correspondence of location of clusters in mediolateral oblique (MLO) and cranio-caudal (CC) views, was used in feature calculation and in final classification. Initially, microcalcifications were automatically detected by using a statistical method based on Bayesian techniques and a Markov random field model. To determine malignancy or benignancy of a cluster, a method based on two classification steps was developed. In the first step, classification of clusters was performed and in the second step a patient based classification was done. A total of 16 features was used in the study. To identify meaningful features, a feature selection was applied, using the area under the receiver operating characteristic (ROC) curve (Az value) as a criterion. For classification the k-nearest-neighbor method was used in a leave-one-patient-out procedure. A database of 192 mammograms with 280 true positive detected microcalcification clusters was used for evaluation of the method. The set consisted of cases that were selected for diagnostic work up during a 4 year period of screening in the Nijmegen region (The Netherlands). Because of the high positive predictive value in the screening program (50%), this set did not contain obvious benign cases. The method's best patient-based performance on this set corresponded with Az = 0.83, using nine features. A subset of the data set, containing mammograms from 90 patients, was used for comparing the computer results to radiologists' performance. Ten radiologists read these cases on a light-box and assessed the probability of malignancy for each patient. All participants had experience in clinical mammography and participated in our observer study during the last 2 days of a 2-week training session leading to screening mammography certification. Results on the subset showed that the method's performance (Az = 0.83) was considerably higher than that of the radiologists (Az = 0.63).

Breast Neoplasms↗

Comparing methods for calculating confidence intervals for vaccine efficacy.

A method is introduced for computing a Bayesian 95 per cent posterior probability region for vaccine efficacy. This method assumes independent vague gamma prior distributions for the incidence rates on each arm of the trial, and a Poisson likelihood for the counts of incident cases of infection. The approach is similar in spirit to the Bayesian analysis of the binomial risk ratio described by Aitchison and Bacon-Shone. However, the focus of our interest is not on incorporating prior information into the design of trials for efficacy, but rather on evaluating whether or not the Bayesian approach with vague prior information produces comparable results to a frequentist approach. A review of methods for constructing exact and large sample intervals for vaccine efficacy is provided as a framework for comparison. The confidence interval methods are assessed by comparing the size and power of tests of vaccine efficacy in proposed intermediate sized randomized double blinded placebo controlled trials.

AIDS Vaccines↗

Are normative expert systems appropriate for diagnostic pathology?

Conceptual models for diagnostic reasoning proposed in the medical literature are presented to stimulate discussion about the issue of the appropriateness of probabilistic knowledge-based systems for medical diagnosis. Evidence is presented to corroborate the authors' view that diagnosis is a problem-solving task, rather than a decision-making task. In the authors' opinion, probabilistic reasoning is better suited to situations dealing with choices for clinical intervention, rather than to those dealing with determining the correct diagnosis. A critique is given of a diagnostic Bayesian expert system for lymph node pathology. In empirical studies, diagnostic Bayesian systems have been shown to typically list the correct diagnosis as the program's first choice 60% to 70% of the time. One reason for this undistinguished level of diagnostic performance is that Bayesian systems are not designed to represent and use knowledge the same way that an expert does.

Bayes Theorem↗

A likelihood approach to calculating risk support intervals.

Genetic risks are usually computed under the assumption that genetic parameters, such as the recombination fraction, are known without error. Uncertainty in the estimates of these parameters must translate into uncertainty regarding the risk. To allow for uncertainties in parameter values, one may employ Bayesian techniques or, in a maximum-likelihood framework, construct a support interval (SI) for the risk. Here we have implemented the latter approach. The SI for the risk is based on the SIs of parameters involved in the pedigree likelihood. As an empirical example, the SI for the risk was calculated for probands who are members of chronic spinal muscular atrophy kindreds. In order to evaluate the accuracy of a risk in genetic counseling situations, we advocate that, in addition to a point estimate, an SI for the risk should be calculated.

Female↗

A two-stage classifier for identification of protein-protein interface residues.

MOTIVATION: The ability to identify protein-protein interaction sites and to detect specific amino acid residues that contribute to the specificity and affinity of protein interactions has important implications for problems ranging from rational drug design to analysis of metabolic and signal transduction networks. RESULTS: We have developed a two-stage method consisting of a support vector machine (SVM) and a Bayesian classifier for predicting surface residues of a protein that participate in protein-protein interactions. This approach exploits the fact that interface residues tend to form clusters in the primary amino acid sequence. Our results show that the proposed two-stage classifier outperforms previously published sequence-based methods for predicting interface residues. We also present results obtained using the two-stage classifier on an independent test set of seven CAPRI (Critical Assessment of PRedicted Interactions) targets. The success of the predictions is validated by examining the predictions in the context of the three-dimensional structures of protein complexes.

Amino Acid Sequence↗

MRI prior computation and parallel tempering algorithm: a probabilistic resolution of the MEG/EEG inverse problem.

Since the MEG inverse problem is ill-posed and admits many possible solutions, it is not possible to give it a single "true" answer. Therefore, we propose here to use a specific probabilistic algorithm to map the full probability distribution of the MEG sources with Markov Chain Monte Carlo methods. Using a Bayesian approach, the probability of the MEG solutions is expressed as the product of the likelihood by the prior probability. To compute the prior and constrain the MEG inverse problem resolution, MRI data are also acquired and automatically processed to determine the brain position and volume. We then use Parallel Tempering algorithm to estimate the full posterior probability and determine the likely solutions of the inverse problem. We illustrate the method with results obtained from the analysis of somatosensory data. This illustrates both the MRI processing for the prior computation, and how the knowledge of the full posterior probability distribution can be used to estimate the position of the sources, as well as their likely extension.

Adult↗