Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Perception viewed as an inverse problem.

The modern study of perception began when Fechner published his 'Elements of Psychophysics' in 1860. This book has guided most perception research ever since. It has become increasingly clear that there are problems with Fechner's approach, which assumes that the percept is completely determined by the sensory input. Fechner's approach cannot explain the processes that allow our percepts to be veridical. Post-Fechnerian schools (Helmholtzian, Structural, Gestalt and Gibsonian) have tried to deal with this problem, but have not been successful. An alternative to the Fechnerian approach is required. This paper describes an alternative that has been developing over the last 20 years within the computer vision community. It treats perceptual interpretation as a solution of an inverse problem that depends critically on the operation of a priori constraints. Contemporary research, which adopted this approach, has concentrated on verifying the usefulness of Bayesian and standard regularization methods. This paper takes the next step; it discusses theoretical and empirical aspects of studying human perception as an inverse problem. It reviews the literature that illustrates the power of the inverse problem approach. This review leads to the suggestion that progress in the study of perception will benefit if the inverse approach were to be adopted by experimentalists, as well as by the computational modelers, who have been actively exploring its potential to date.

Bayes Theorem↗

Bayesian analysis of haplotypes for linkage disequilibrium mapping.

Haplotype analysis of disease chromosomes can help identify probable historical recombination events and localize disease mutations. Most available analyses use only marginal and pairwise allele frequency information. We have developed a Bayesian framework that utilizes full haplotype information to overcome various complications such as multiple founders, unphased chromosomes, data contamination, and incomplete marker data. A stochastic model is used to describe the dependence structure among several variables characterizing the observed haplotypes, for example, the ancestral haplotypes and their ages, mutation rate, recombination events, and the location of the disease mutation. An efficient Markov chain Monte Carlo algorithm was developed for computing the estimates of the quantities of interest. The method is shown to perform well in both real data sets (cystic fibrosis data and Friedreich ataxia data) and simulated data sets. The program that implements the proposed method, BLADE, as well as the two real datasets, can be obtained from http://www.fas.harvard.edu/~junliu/TechRept/01folder/diseq_prog.tar.gz.

Bayes Theorem↗

Selecting the best-fit model of nucleotide substitution.

Despite the relevant role of models of nucleotide substitution in phylogenetics, choosing among different models remains a problem. Several statistical methods for selecting the model that best fits the data at hand have been proposed, but their absolute and relative performance has not yet been characterized. In this study, we compare under various conditions the performance of different hierarchical and dynamic likelihood ratio tests, and of Akaike and Bayesian information methods, for selecting best-fit models of nucleotide substitution. We specifically examine the role of the topology used to estimate the likelihood of the different models and the importance of the order in which hypotheses are tested. We do this by simulating DNA sequences under a known model of nucleotide substitution and recording how often this true model is recovered by the different methods. Our results suggest that model selection is reasonably accurate and indicate that some likelihood ratio test methods perform overall better than the Akaike or Bayesian information criteria. The tree used to estimate the likelihood scores does not influence model selection unless it is a randomly chosen tree. The order in which hypotheses are tested, and the complexity of the initial model in the sequence of tests, influence model selection in some cases. Model fitting in phylogenetics has been suggested for many years, yet many authors still arbitrarily choose their models, often using the default models implemented in standard computer programs for phylogenetic estimation. We show here that a best-fit model can be readily identified. Consequently, given the relevance of models, model fitting should be routine in any phylogenetic analysis that uses models of evolution.

Algorithms↗

A Bayesian approach to logistic regression models having measurement error following a mixture distribution.

To estimate the parameters in a logistic regression model when the predictors are subject to random or systematic measurement error, we take a Bayesian approach and average the true logistic probability over the conditional posterior distribution of the true value of the predictor given its observed value. We allow this posterior distribution to consist of a mixture when the measurement error distribution changes form with observed exposure. We apply the method to study the risk of alcohol consumption on breast cancer using the Nurses Health Study data. We estimate measurement error from a small subsample where we compare true with reported consumption. Some of the self-reported non-drinkers truly do not drink. The resulting risk estimates differ sharply from those computed by standard logistic regression that ignores measurement error.

Age Factors↗

Uncertainties in doses from intakes of radionuclides assessed from monitoring measurements.

The evaluation of uncertainties in doses from intakes of radionuclides is one of the most difficult problems in internal dosimetry. In this paper, the process of assessing internal doses from monitoring measurements is reviewed and the major sources of uncertainty are discussed. Methods developed independently at HPA and at IRSN for the determination of uncertainties in internal doses assessed from monitoring measurements are described. Both use a Monte Carlo simulation approach. Results are described for three illustrative examples. An alternative method developed at the Los Alamos National Laboratory that uses Bayesian statistical methods is also described briefly.

Animals↗

An example of complex modelling in dentistry using Markov chain Monte Carlo (MCMC) simulation.

BACKGROUND: In the usual regression setting one regression line is computed for a whole data set. In a more complex situation, each person may be observed for example at several points in time and thus a regression line might be calculated for each person. Additional complexities, such as various forms of errors in covariables may make a straightforward statistical evaluation difficult or even impossible. OBJECTIVE AND METHOD: During recent years methods have been developed allowing convenient analysis of problems where the data and the corresponding models show these and many other forms of complexity. The methodology makes use of a Bayesian approach and Markov chain Monte Carlo (MCMC) simulations. The methods allow the construction of increasingly elaborate models by building them up from local sub-models. The essential structure of the models can be represented visually by directed acyclic graphs (DAG). This attractive property allows communication and discussion of the essential structure and the substantial meaning of a complex model without needing algebra. EXAMPLE: After presentation of the statistical methods an example from dentistry is presented in order to demonstrate their application and use. The dataset of the example had a complex structure; each of a set of children was followed up over several years. The number of new fillings in permanent teeth had been recorded at several ages. The dependent variables were markedly different from the normal distribution and could not be transformed to normality. In addition, explanatory variables were assumed to be measured with different forms of error. Illustration of how the corresponding models can be estimated conveniently via MCMC simulation, in particular, 'Gibbs sampling', using the freely available software BUGS is presented. In addition, how the measurement error may influence the estimates of the corresponding coefficients is explored. It is demonstrated that the effect of the independent variable on the dependent variable may be markedly underestimated if the measurement error is not taken into account ('regression dilution bias'). CONCLUSION: Markov chain Monte Carlo methods may be of great value to dentists in allowing analysis of data sets which exhibit a wide range of different forms of complexity.

Bayes Theorem↗

Study and performance evaluation of statistical methods in image processing.

Two statistical image processing formalisms involving the entropy concept and Bayesian analysis are studied. Iterative imaging algorithms of the formalisms are formulated by employing, for the purpose of performance evaluation and easy implementation, the steepest descent method for the solution of entropy concept and the expectation maximization technique for the solution of Bayesian analysis. Quantitative evaluation and comparison of the convergence performance of the iterative algorithms on computer generated ideal and experimental radioisotope phantom imaging noisy data are given. The study concludes that the entropy algorithm can converge relatively fast, but it is very sensitive to noise in measured data due to the ill-posed nature of inverse problems and its lack of ability to consider the statistics of data fluctuation; while the Bayesian algorithm converges monotonically even with noisy data and has the advantage of considering both the a priori source distribution information and the statistical fluctuation of measured data.

Algorithms↗

Predicting essential genes in fungal genomes.

Essential genes are required for an organism's viability, and the ability to identify these genes in pathogens is crucial to directed drug development. Predicting essential genes through computational methods is appealing because it circumvents expensive and difficult experimental screens. Most such prediction is based on homology mapping to experimentally verified essential genes in model organisms. We present here a different approach, one that relies exclusively on sequence features of a gene to estimate essentiality and offers a promising way to identify essential genes in unstudied or uncultured organisms. We identified 14 characteristic sequence features potentially associated with essentiality, such as localization signals, codon adaptation, GC content, and overall hydrophobicity. Using the well-characterized baker's yeast Saccharomyces cerevisiae, we employed a simple Bayesian framework to measure the correlation of each of these features with essentiality. We then employed the 14 features to learn the parameters of a machine learning classifier capable of predicting essential genes. We trained our classifier on known essential genes in S. cerevisiae and applied it to the closely related and relatively unstudied yeast Saccharomyces mikatae. We assessed predictive success in two ways: First, we compared all of our predictions with those generated by homology mapping between these two species. Second, we verified a subset of our predictions with eight in vivo knockouts in S. mikatae, and we present here the first experimentally confirmed essential genes in this species.

Computational Biology↗

Understanding tuberculosis epidemiology using structured statistical models.

Molecular epidemiological studies can provide novel insights into the transmission of infectious diseases such as tuberculosis. Typically, risk factors for transmission are identified using traditional hypothesis-driven statistical methods such as logistic regression. However, limitations become apparent in these approaches as the scope of these studies expand to include additional epidemiological and bacterial genomic data. Here we examine the use of Bayesian models to analyze tuberculosis epidemiology. We begin by exploring the use of Bayesian networks (BNs) to identify the distribution of tuberculosis patient attributes (including demographic and clinical attributes). Using existing algorithms for constructing BNs from observational data, we learned a BN from data about tuberculosis patients collected in San Francisco from 1991 to 1999. We verified that the resulting probabilistic models did in fact capture known statistical relationships. Next, we examine the use of newly introduced methods for representing and automatically constructing probabilistic models in structured domains. We use statistical relational models (SRMs) to model distributions over relational domains. SRMs are ideally suited to richly structured epidemiological data. We use a data-driven method to construct a statistical relational model directly from data stored in a relational database. The resulting model reveals the relationships between variables in the data and describes their distribution. We applied this procedure to the data on tuberculosis patients in San Francisco from 1991 to 1999, their Mycobacterium tuberculosis strains, and data on contact investigations. The resulting statistical relational model corroborated previously reported findings and revealed several novel associations. These models illustrate the potential for this approach to reveal relationships within richly structured data that may not be apparent using conventional statistical approaches. We show that Bayesian methods, in particular statistical relational models, are an important tool for understanding infectious disease epidemiology.

Adult↗

Pharmacovigilance in the 21st century: new systematic tools for an old problem.

The large number of adverse-event reports generated by marketed drugs and devices argues for the application of validated computerized algorithms to supplement traditional methods of detecting adverse-event signals. Difficulties in accurately estimating patient exposure and background rates for a given event in a specific population hinder risk estimation in spontaneous adverse-event databases. The United States Food and Drug Administration (FDA) is evaluating a Bayesian data mining system called Multi-item Gamma Poisson Shrinker (MGPS) to enhance the FDA's ability to monitor the safety of drugs, biologics, and vaccines after they have been approved for use. The MGPS computes adjusted higher-than-expected reporting relationships between drugs and adverse events across 35 years of data relative to internal background rates. The MGPS can also adjust for random noise by using a model derived from the data, and corrects for temporal trends and confounding related to age, sex, and other variables by stratifying over 900 categories. Signals can then be compared with or used in conjunction with other sources (e.g. clinical trials, general practice databases) to further study the adverse-event risk. The example of pancreatitis risk with atypical antipsychotics, valproic acid, and valproate is used to discuss the strengths and limitations of MGPS versus traditional methods. Validated data mining techniques offer great promise to enhance pharmacovigilance practices.

Adverse Drug Reaction Reporting Systems↗

Population toxicokinetic analysis of 2,3,7,8-tetrachlorodibenzo-p-dioxin using Bayesian techniques.

Understanding the kinetics of 2,3,7,8-tetrachlorodibenzo-p-dioxin (TCDD) concentrations in humans is an important step for TCDD cancer risk assessment. In this paper longitudinal series of serum TCDD concentration measurements on U.S. veterans of the Vietnam war, who were exposed to dioxin during herbicide-spraying operations, are studied. The overall aim is to use these data to infer the dynamics of TCDD concentrations in humans. This is done by identifying a kinetic model describing the dioxin time course at the individual level. The individual toxicokinetic model is then expanded into a population model within a Bayesian hierarchical framework which allows residual variations across subjects that cannot be explained by observed covariates. Other complications in the data, such as unknown exposure histories, are also resolved implicitly through the hierarchical model. Moreover, the choice of a Bayesian approach enables the accumulation of external source of information in the form of prior distributions. The model is subjected to various diagnostic checks and analyses of sensitivity to distributional assumptions showing a good fit in terms of both the population and the kinetic features.

Bayes Theorem↗

Use of gene networks for identifying and validating drug targets.

We propose a new method for identifying and validating drug targets by using gene networks, which are estimated from cDNA microarray gene expression profile data. We created novel gene disruption and drug response microarray gene expression profile data libraries for the purpose of drug target elucidation. We use two types of microarray gene expression profile data for estimating gene networks and then identifying drug targets. The estimated gene networks play an essential role in understanding drug response data and this information is unattainable from clustering methods, which are the standard for gene expression analysis. In the construction of gene networks, we use the Bayesian network model. We use an actual example from analysis of the Saccharomyces cerevisiae gene expression profile data to express a concrete strategy for the application of gene network information to drug discovery.

Algorithms↗

Exploring the relationship between rationality and bounded rationality in medical knowledge-based systems.

If our goal in Artificial Intelligence in Medicine (AIM) is to engineer systems health-care providers will both use and, in the process, improve their performance, we must concentrate on the development of causal theories of knowledge and problem solving. One broad direction in pursuing this goal is understanding the relationships between existing models of rationality and bounded rationality for similar tasks. Models of rationality refer to those approaches in which the optimal properties of the models are deductively provable, i.e. in which the processing is rational. Representative models of rationality used in AIM are deductive logical models, statistical models such as Bayesian inference models, and decision-analytic models. Models of bounded rationality are those which do not guarantee such optimal properties nor yield to deductive correctness proofs. These models have their roots in cognitive psychology. In this article we show how explicating the relationship between models of rationality and bounded rationality might be done in the case of abductive tasks in medicine. This is done by positioning these modeling approaches within the same framework (an abstract computational model) and interpreting in this context both computational complexity results concerning the nature of the task and empirical results studies of human problem-solving behavior.

Artificial Intelligence↗

A quantitative structure--activity relationships model for the acute toxicity of substituted benzenes to Tetrahymena pyriformis using Bayesian-regularized neural networks.

We have used a new, robust structure-activity mapping technique, a Bayesian-regularized neural network, to develop a quantitative structure-activity relationships (QSAR) model for the toxicity of 278 substituted benzenes toward Tetrahymena pyriformis. The independent variables used in the modeling were derived solely from the molecular structure, and the model was tested on 20% of the data set selected from the whole set by cluster analysis and which had not been used in training the network. The results show that the method is robust and reliable and give results for mixed class compounds which are comparable to earlier QSAR work on single-chemical class subsets of the 278 compounds and which employed measured physicochemical parameters as independent variables. Comparisons of Bayesian neural net models with those derived by classical PLS analysis showed the superiority of our method. The method appears to be able to model more diverse chemical classes and more than one mechanism of toxicity.

Animals↗

Range image segmentation by an effective jump-diffusion method.

This paper presents an effective jump-diffusion method for segmenting a range image and its associated reflectance image in the Bayesian framework. The algorithm works on complex real-world scenes (indoor and outdoor), which consist of an unknown number of objects (or surfaces) of various sizes and types, such as planes, conics, smooth surfaces, and cluttered objects (like trees and bushes). Formulated in the Bayesian framework, the posterior probability is distributed over a solution space with a countable number of subspaces of varying dimensions. The algorithm simulates Markov chains with both reversible jumps and stochastic diffusions to traverse the solution space. The reversible jumps realize the moves between subspaces of different dimensions, such as switching surface models and changing the number of objects. The stochastic Langevin equation realizes diffusions within each subspace. To achieve effective computation, the algorithm precomputes some importance proposal probabilities over multiple scales through Hough transforms, edge detection, and data clustering. The latter are used by the Markov chains for fast mixing. The algorithm is tested on 100 1D simulated data sets for performance analysis on both accuracy and speed. Then, the algorithm is applied to three data sets of range images under the same parameter setting. The results are satisfactory in comparison with manual segmentations.

Algorithms↗

Selecting factors predictive of heterogeneity in multivariate event time data.

In multivariate survival analysis, investigators are often interested in testing for heterogeneity among clusters, both overall and within specific classes. We represent different hypotheses about the heterogeneity structure using a sequence of gamma frailty models, ranging from a null model with no random effects to a full model having random effects for each class. Following a Bayesian approach, we define prior distributions for the frailty variances consisting of mixtures of point masses at zero and inverse-gamma densities. Since frailties with zero variance effectively drop out of the model, this prior allocates probability to each model in the sequence, including the overall null hypothesis of homogeneity. Using a counting process formulation, the conditional posterior distributions of the frailties and proportional hazards regression coefficients have simple forms. Posterior computation proceeds via a data augmentation Gibbs sampling algorithm, a single run of which can be used to obtain model-averaged estimates of the population parameters and posterior model probabilities for testing hypotheses about the heterogeneity structure. The methods are illustrated using data from a lung cancer trial.

Algorithms↗

A Bayesian approach to introducing anatomo-functional priors in the EEG/MEG inverse problem.

In this paper, we present a new approach to the recovering of dipole magnitudes in a distributed source model for magnetoencephalographic (MEG) and electroencephalographic (EEG) imaging. This method consists in introducing spatial and temporal a priori information as a cure to this ill-posed inverse problem. A nonlinear spatial regularization scheme allows the preservation of dipole moment discontinuities between some a priori noncorrelated sources, for instance, when considering dipoles located on both sides of a sulcus. Moreover, we introduce temporal smoothness constraints on dipole magnitude evolution, at time scales smaller than those of cognitive processes. These priors are easily integrated into a Bayesian formalism, yielding a maximum a posteriori (MAP) estimator of brain electrical activity. Results from EEG simulations of our method are presented and compared with those of classical quadratic regularization and a now popular generalized minimum-norm technique called low-resolution electromagnetic tomography (LORETA).

Algorithms↗

A bayesian hierarchical model for the analysis of Affymetrix arrays.

An area of active research in DNA microarray analysis focuses on identifying differentially expressed genes between normal and malignant tissues. The analysis is complicated by the presence of several unreliable expression readings. Here, we illustrate a methodology where the expression estimates are modeled as censored data and discriminating genes are selected using ANOVA-based criteria.

Analysis of Variance↗