Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Classification by multiple-resolution statistical analysis with application to automated recognition of marine mammal sounds.

A multiple-resolution statistical pattern recognition technique for classification by supervised learning is developed and then applied to automated recognition of marine mammal sounds. The data to be classified may be either unprocessed or transformed, e.g., time series or time-frequency distributions of acoustic transients. Training data consist of samples previously grouped by a human expert into labeled sets; these sets are presumed to be associated with different "classes." The labeled sets are then characterized by occupancy statistics associated with a multiple-resolution, binary partition of the (unreduced) sample space. Classification of a new sample is performed by calculating a posteriori probabilities of membership of the new sample in each class, computed by Bayesian inference from the occupancy statistics of the associated labeled set. These a posteriori probabilities are calculated by a recursive algorithm that progresses from coarse to fine resolution in the sample space. The algorithm is implemented in a simple, highly efficient computer program. Automated classification of both time series and time-frequency distributions of marine-mammal vocalizations is demonstrated using a small number of labeled samples (approximately ten samples per class).

Algorithms↗

A study of early stopping and model selection applied to the papermaking industry.

This paper addresses the issues of neural network model development and maintenance in the context of a complex task taken from the papermaking industry. In particular, it describes a comparison study of early stopping techniques and model selection, both to optimise neural network models for generalisation performance. The results presented here show that early stopping via use of a Bayesian model evidence measure is a viable way of optimising performance while also making maximum use of all the data. In addition, they show that ten-fold cross-validation performs well as a model selector and as an estimator of prediction accuracy. These results are important in that they show how neural network models may be optimally trained and selected for highly complex industrial tasks where the data are noisy and limited in number.

Algorithms↗

Bayesian population pharmacokinetic and pharmacodynamic analyses using mixture models.

Population studies of the pharmacokinetics or pharmacodynamics of drugs help us, learn about the variability in drug disposition and effects, information that can be used to treat future patients at safe and effective doses. We present a new approach to population modeling based on a weighted mixture of normal distributions having random weights and means. This method allows estimation of underlying continuous population distributions without prespecifying the parametric form or shape of these probability distributions. Additionally, this method can carry out nonparametric regression of pharmacokinetic or dynamic parameters on patient covariates while estimating the underlying distributions. Two examples illustrate the method and its flexibility.

Bayes Theorem↗

Exploring heterogeneity in tumour data using Markov chain Monte Carlo.

We describe a Bayesian approach to incorporate between-individual heterogeneity associated with parameters of complicated biological models. We emphasize the use of the Markov chain Monte Carlo (MCMC) method in this context and demonstrate the implementation and use of MCMC by analysis of simulated overdispersed Poisson counts and by analysis of an experimental data set on preneoplastic liver lesions (their number and sizes) in the presence of heterogeneity. These examples show that MCMC-based estimates, derived from the posterior distribution with uniform priors, may agree well with maximum likelihood estimates (if available). However, with heterogeneous parameters, maximum likelihood estimates can be difficult to obtain, involving many integrations. In this case, the MCMC method offers substantial computational advantages.

Animals↗

Bayesian cost-effectiveness analysis from clinical trial data.

A key tool for assessing the relative cost-effectiveness of two treatments in health economics is the incremental C/E acceptability curve. We present Bayesian computations for this curve in the case where data on both costs and efficacy are available from a clinical trial. Analysis is given under various formulations of prior information. A case study is analysed in which reasonable prior information is shown to strengthen substantially the posterior inference, leading to a more conclusive assessment of cost-effectiveness. Calculations can be performed using readily available Bayesian software.

Anti-Asthmatic Agents↗

Backtracking booze with Bayes--the retrospective interpretation of blood alcohol data.

1. A Bayesian method is described which allows the explicit estimation of errors produced in estimating drug concentrations at times for which samples are not available for analysis. 2. This method was applied to the problem of 'backtracking' alcohol concentrations for medico-legal purposes. 3. Computer simulation allowed the effect of continuing alcohol absorption on the position and range of estimates of alcohol concentrations to be studied.

Accidents↗

Mining genetic epidemiology data with Bayesian networks I: Bayesian networks and example application (plasma apoE levels).

MOTIVATION: The wealth of single nucleotide polymorphism (SNP) data within candidate genes and anticipated across the genome poses enormous analytical problems for studies of genotype-to-phenotype relationships, and modern data mining methods may be particularly well suited to meet the swelling challenges. In this paper, we introduce the method of Belief (Bayesian) networks to the domain of genotype-to-phenotype analyses and provide an example application. RESULTS: A Belief network is a graphical model of a probabilistic nature that represents a joint multivariate probability distribution and reflects conditional independences between variables. Given the data, optimal network topology can be estimated with the assistance of heuristic search algorithms and scoring criteria. Statistical significance of edge strengths can be evaluated using Bayesian methods and bootstrapping. As an example application, the method of Belief networks was applied to 20 SNPs in the apolipoprotein (apo) E gene and plasma apoE levels in a sample of 702 individuals from Jackson, MS. Plasma apoE level was the primary target variable. These analyses indicate that the edge between SNP 4075, coding for the well-known epsilon2 allele, and plasma apoE level was strong. Belief networks can effectively describe complex uncertain processes and can both learn from data and incorporate prior knowledge. AVAILABILITY: Various alternative and supplemental networks (not given in the text) as well as source code extensions, are available from the authors. SUPPLEMENTARY INFORMATION: http://bioinformatics.oxfordjournals.org.

Apolipoproteins E↗

Bayesian image processing in magnetic resonance imaging.

In the past several years, image processing techniques based on Bayesian models have received considerable attention. In our earlier work, we developed a novel Bayesian approach which was primarily aimed at the processing and reconstruction of images in positron emission tomography. In this paper, we describe how the technique has been adopted to process magnetic resonance images in order to reduce noise and artifacts, thereby improving image quality. In this framework, the image is assumed to be a statistical variable whose posterior probability density conditional on the observed image is modeled by the product of the likelihood function of the observed data with a prior density based our prior knowledge. A Gibbs random field incorporating local continuity information and with edge-detection capability is used as the prior model. Based on the formalism of the posterior density, we can compute an estimate of the image using an iterative technique. We have implemented this technique and applied it to phantom and clinical images. Our results indicate that the approach works reasonably well for reducing noise, enhancing edges, and removing ringing artifact.

Algorithms↗

Uncertainty, neuromodulation, and attention.

Uncertainty in various forms plagues our interactions with the environment. In a Bayesian statistical framework, optimal inference and prediction, based on unreliable observations in changing contexts, require the representation and manipulation of different forms of uncertainty. We propose that the neuromodulators acetylcholine and norepinephrine play a major role in the brain's implementation of these uncertainty computations. Acetylcholine signals expected uncertainty, coming from known unreliability of predictive cues within a context. Norepinephrine signals unexpected uncertainty, as when unsignaled context switches produce strongly unexpected observations. These uncertainty signals interact to enable optimal inference and learning in noisy and changeable environments. This formulation is consistent with a wealth of physiological, pharmacological, and behavioral data implicating acetylcholine and norepinephrine in specific aspects of a range of cognitive processes. Moreover, the model suggests a class of attentional cueing tasks that involve both neuromodulators and shows how their interactions may be part-antagonistic, part-synergistic.

Acetylcholine↗

Building statistical models to analyze species distributions.

Models of the geographic distributions of species have wide application in ecology. But the nonspatial, single-level, regression models that ecologists have often employed do not deal with problems of irregular sampling intensity or spatial dependence, and do not adequately quantify uncertainty. We show here how to build statistical models that can handle these features of spatial prediction and provide richer, more powerful inference about species niche relations, distributions, and the effects of human disturbance. We begin with a familiar generalized linear model and build in additional features, including spatial random effects and hierarchical levels. Since these models are fully specified statistical models, we show that it is possible to add complexity without sacrificing interpretability. This step-by-step approach, together with attached code that implements a simple, spatially explicit, regression model, is structured to facilitate self-teaching. All models are developed in a Bayesian framework. We assess the performance of the models by using them to predict the distributions of two plant species (Proteaceae) from South Africa's Cape Floristic Region. We demonstrate that making distribution models spatially explicit can be essential for accurately characterizing the environmental response of species, predicting their probability of occurrence, and assessing uncertainty in the model results. Adding hierarchical levels to the models has further advantages in allowing human transformation of the landscape to be taken into account, as well as additional features of the sampling process.

Bayes Theorem↗

Bayesian modeling of the hemodynamic response function in BOLD fMRI.

In functional magnetic resonance imaging (fMRI), modeling the complex link between neuronal activity and its hemodynamic response via the neurovascular coupling requires an elaborate and sensitive response model. Methods based on physiologic assumptions as well as direct, descriptive models have been proposed. The focus of this study is placed on such a direct approach that allows for a robust pixelwise determination of hemodynamic characteristics, such as time to peak or the poststimulus undershoot. A Bayesian procedure is presented that can easily be adapted to different hemodynamic properties in question and can be estimated without numerical problems known from nonlinear optimization algorithms. The usefulness of the model is demonstrated by thorough analyzes of the poststimulus undershoot in visual and acoustic stimulation paradigms. Further, we show the capability of this approach to improve analysis of fMRI data in altered hemodynamic conditions.

Auditory Perception↗

A classical likelihood based approach for admixture mapping using EM algorithm.

Several disease-mapping methods have been proposed recently, which use the information generated by recent admixture of populations from historically distinct geographic origins. These methods include both classic likelihood and Bayesian approaches. In this study we directly maximize the likelihood function from the hidden Markov Model for admixture mapping using the EM algorithm, allowing for uncertainty in model parameters, such as the allele frequencies in the parental populations. We determined the robustness of the proposed method by examining the ancestral allele frequency estimate and individual marker-location specific ancestry when the data were generated by different population admixture models and no learning sample was used. The proposed method outperforms a widely used Bayesian MCMC strategy for data generated from various population admixture models. The multipoint information content for ancestry was derived based on the map provided by Smith et al. (2004) and the associated statistical power was calculated. We examined the distribution of admixture LD across the genome for both real and simulated data and established a threshold for genome wide significance applicable to admixture mapping studies. The software ADMIXPROGRAM for performing admixture mapping is available from authors.

Algorithms↗

Stochastic complexities of reduced rank regression in Bayesian estimation.

Reduced rank regression extracts an essential information from examples of input-output pairs. It is understood as a three-layer neural network with linear hidden units. However, reduced rank approximation is a non-regular statistical model which has a degenerate Fisher information matrix. Its generalization error had been left unknown even in statistics. In this paper, we give the exact asymptotic form of its generalization error in Bayesian estimation, based on resolution of learning machine singularities. For this purpose, the maximum pole of the zeta function for the learning theory is calculated. We propose a new method of recursive blowing-ups which yields the complete desingularization of the reduced rank approximation.

Artificial Intelligence↗

Dependence among sites in RNA evolution.

Although probabilistic models of genotype (e.g., DNA sequence) evolution have been greatly elaborated, less attention has been paid to the effect of phenotype on the evolution of the genotype. Here we propose an evolutionary model and a Bayesian inference procedure that are aimed at filling this gap. In the model, RNA secondary structure links genotype and phenotype by treating the approximate free energy of a sequence folded into a secondary structure as a surrogate for fitness. The underlying idea is that a nucleotide substitution resulting in a more stable secondary structure should have a higher rate than a substitution that yields a less stable secondary structure. This free energy approach incorporates evolutionary dependencies among sequence positions beyond those that are reflected simply by jointly modeling change at paired positions in an RNA helix. Although there is not a formal requirement with this approach that secondary structure be known and nearly invariant over evolutionary time, computational considerations make these assumptions attractive and they have been adopted in a software program that permits statistical analysis of multiple homologous sequences that are related via a known phylogenetic tree topology. Analyses of 5S ribosomal RNA sequences are presented to illustrate and quantify the strong impact that RNA secondary structure has on substitution rates. Analyses on simulated sequences show that the new inference procedure has reasonable statistical properties. Potential applications of this procedure, including improved ancestral sequence inference and location of functionally interesting sites, are discussed.

Animals↗

Modeling gene expression from microarray expression data with state-space equations.

We describe a new method to model gene expression from time-course gene expression data. The modelling is in terms of state-space descriptions of linear systems. A cell can be considered to be a system where the behaviours (responses) of the cell depend completely on the current internal state plus any external inputs. The gene expression levels in the cell provide information about the behaviours of the cell. In previously proposed methods, genes were viewed as internal state variables of a cellular system and their expression levels were the values of the intemal state variables. This viewpoint has suffered from the underestimation of the model parameters. Instead, we view genes as the observation variables, whose expression values depend on the current intemal state variables and any external input. Factor analysis is used to identify the internal state variables, and Bayesian Information Criterion (BIC) is used to determine the number of the internal state variables. By building dynamic equations of the internal state variables and the relationships between the internal state variables and the observation variables (gene expression profiles), we get state-space descriptions of gene expression model. In the present method, model parameters may be unambiguously identified from time-course gene expression data. We apply the method to two time-course gene expression datasets to illustrate it.

Algorithms↗

Estimating the probability of the presence of a signal of interest in multiresolution single- and multiband image denoising.

We develop three novel wavelet domain denoising methods for subband-adaptive, spatially-adaptive and multivalued image denoising. The core of our approach is the estimation of the probability that a given coefficient contains a significant noise-free component, which we call "signal of interest." In this respect, we analyze cases where the probability of signal presence is 1) fixed per subband, 2) conditioned on a local spatial context, and 3) conditioned on information from multiple image bands. All the probabilities are estimated assuming a generalized Laplacian prior for noise-free subband data and additive white Gaussian noise. The results demonstrate that the new subband-adaptive shrinkage function outperforms Bayesian thresholding approaches in terms of mean-squared error. The spatially adaptive version of the proposed method yields better results than the existing spatially adaptive ones of similar and higher complexity. The performance on color and on multispectral images is superior with respect to recent multiband wavelet thresholding.

Algorithms↗

[Automatic diagnosis of malignant degree of brain glioma based on Bayesian network].

Bayesian network connects graph theory with statistics, being an important research direction in data mining. Compared with other approaches used for data mining, Bayesian network can combine prior knowledge with observed data. Besides that, it can handle incomplete data sets. This paper applies Bayesian network to predict the malignant degree of brain glioma. Totally 280 cases are collected, and some of them contain missing values. Preprocessing is taken to make them applicable to the algorithms. Unlike MLP network, both Bayesian network and decision tree use attribute-value pairs to represent diagnostic knowledge derived from treated cases. These could improve both the understandability and applicability of their results. Results of all these algorithms can achieve accuracy rate over 80%, which satisfies the requirement of neuroradiologists.

Bayes Theorem↗

Implications of unchanged detection criteria with CAD as second reader of mammograms.

In this paper we address the use of computer-aided detection (CAD) systems as second readers in mammography. The approach is based on Bayesian decision theory and its implication for the choice of optimal operating points. The choice of a certain operating point along an ROC curve corresponds to a particular tradeoff between false positives and missed cancers. By minimizing a total risk function given this tradeoff, we determine optimal decision thresholds for the radiologist and CAD system when CAD is used as a second reader. We show that under very general circumstances, the performance of the sequential system is improved if the decision threshold of the latent human decision variable is increased compared to what it would have been in the absence of the CAD system. This means that an initial stricter decision criterion should be applied by the radiologist when CAD is used as a second reader than otherwise. First and foremost, the results in this paper should be interpreted qualitatively, but an attempt is made at quantifying the effect by tuning the model to a prospective study evaluating the use of CAD as a second reader. By making some necessary and plausible assumptions, we are able to estimate the effect of the resulting suboptimal operating point. In this study of 12 860 women, we estimate that a 15% reduction in callbacks for masses could have been achieved with only about a 1.5% relative decrease in sensitivity compared to that without using a stricter initial criterion by the radiologist. For microcalcifications the corresponding values are 7% and 0.2%.

Breast Neoplasms↗