Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Generalized conjugate priors for Bayesian analysis of risk and survival regressions.

Conjugate priors for Bayesian analyses of relative risks can be quite restrictive, because their shape depends on their location. By introducing a separate location parameter, however, these priors generalize to allow modeling of a broad range of prior opinions, while still preserving the computational simplicity of conjugate analyses. The present article illustrates the resulting generalized conjugate analyses using examples from case-control studies of the association of residential wire codes and magnetic fields with childhood leukemia.

Bayes Theorem↗

A comparison of methods for estimating the transition:transversion ratio from DNA sequences.

Estimation of the ratio of the rates of transitions to transversions (TI:TV ratio) for a collection of aligned nucleotide sequences is important because it provides insight into the process of molecular evolution and because such estimates may be used to further model the evolutionary process for the sequences under consideration. In this paper, we compare several methods for estimating the TI:TV ratio, including the pairwise method [TREE 11 (1996) 158], a modification of the pairwise method due to Ina [J. Mol. Evol. 46 (1998) 521], a method based on parsimony (TREE 11 (1996) 158), a method due to Purvis and Bromham [J. Mol. Evol. 44 (1997) 112] that uses phylogenetically independent pairs of sequences, the maximum likelihood method, and a Bayesian method [Bioinformatics 17 (2001) 754]. We examine the performance of each estimator under several conditions using both simulated and real data.

Animals↗

Estimating immunoregulatory gene networks in human herpesvirus type 6-infected T cells.

The immune response to viral infection involves complex network of dynamic gene and protein interactions. We present here the dynamic gene network of the host immune response during human herpesvirus type 6 (HHV-6) infection in an adult T-cell leukemia cell line. Using a pathway-focused oligonucleotide DNA microarray, we found a possible association between chemokine genes regulating Th1/Th2 balance and genes regulating T-cell proliferation during HHV-6B infection. Gene network analysis using an integrated comprehensive workbench, VoyaGene, revealed that a gene encoding a TEC-family kinase, ITK, might be a putative modulator in the host immune response against HHV-6B infection. We conclude that Th2-dominated inflammatory reaction in host cells may play an important role in HHV-6B-infected T cells, thereby suggesting the possibility that ITK might be a therapeutic target in diseases related to dysregulation of Th1/Th2 balance. This study describes a novel approach to find genes related with the complex host-virus interaction using microarray data employing the Bayesian statistical framework.

Adult↗

Spike sorting: Bayesian clustering of non-stationary data.

Spike sorting involves clustering spikes recorded by a micro-electrode according to the source neurons. It is a complicated task, which requires much human labor, in part due to the non-stationary nature of the data. We propose to automate the clustering process in a Bayesian framework, with the source neurons modeled as a non-stationary mixture-of-Gaussians. At a first search stage, the data are divided into short time frames, and candidate descriptions of the data as mixtures-of-Gaussians are computed for each frame separately. At a second stage, transition probabilities between candidate mixtures are computed, and a globally optimal clustering solution is found as the maximum-a-posteriori solution of the resulting probabilistic model. The transition probabilities are computed using local stationarity assumptions, and are based on a Gaussian version of the Jensen-Shannon divergence. We employ synthetically generated spike data to illustrate the method and show that it outperforms other spike sorting methods in a non-stationary scenario. We then use real spike data and find high agreement of the method with expert human sorters in two modes of operation: a fully unsupervised and a semi-supervised mode. Thus, this method differs from other methods in two aspects: its ability to account for non-stationary data, and its close to human performance.

Algorithms↗

Multivariate autoregressive modeling of fMRI time series.

We propose the use of multivariate autoregressive (MAR) models of functional magnetic resonance imaging time series to make inferences about functional integration within the human brain. The method is demonstrated with synthetic and real data showing how such models are able to characterize interregional dependence. We extend linear MAR models to accommodate nonlinear interactions to model top-down modulatory processes with bilinear terms. MAR models are time series models and thereby model temporal order within measured brain activity. A further benefit of the MAR approach is that connectivity maps may contain loops, yet exact inference can proceed within a linear framework. Model order selection and parameter estimation are implemented by using Bayesian methods.

Algorithms↗

Experimental measurements of fluence distribution in a UV reactor using fluorescent microspheres.

One concern with current techniques of UV reactor validation is that they provide only a measure of the mean UV fluence. In this research, the actual fluence distribution of a UV reactor is measured through the use of photochemically active fluorescent microspheres. Experimental tests were performed in a pilot-scale monochromatic UV 254 nm reactor operated at two flow rates. Analysis of the fluorescence intensity decay was performed using collimated beam experiments for determination of decay rate kinetics. A stochastic hierarchal process involving Bayesian statistics, and the Markov chain Monte Carlo integration technique was used to correlate the microsphere fluorescence intensity distribution to the UV fluence distribution. The experimental UV fluence distribution was compared with the fluence distribution predicted using a computational fluid dynamics model. The results showed that the fluorescent microspheres measured a wider distribution of UV fluences with a flow rate of 3 gpm than with 7.5 gpm. The principal differences between the modeled and the measured distribution were in the low UV fluences where the microspheres predicted lower fluence levels than the model. The use of microspheres is demonstrated as a novel technique for measurement of the fluence distribution in UV reactors. This technique has both fundamental and practical implications for reactor evaluation and testing and could improve confidence in the future use of mathematical models for UV reactor characterization. It also serves as a complement to biodosimetry testing by providing greater insights regarding reactor behavior and validation.

Bacillus subtilis↗

Accelerated median root prior reconstruction for pinhole single-photon emission tomography (SPET).

Pinhole collimation can be used to improve spatial resolution in SPET. However, the resolution improvement is achieved at the cost of reduced sensitivity, which leads to projection images with poor statistics. Images reconstructed from these projections using the maximum likelihood expectation maximization (ML-EM) algorithms, which have been used to reduce the artefacts generated by the filtered backprojection (FBP) based reconstruction, suffer from noise/bias trade-off: noise contaminates the images at high iteration numbers, whereas early abortion of the algorithm produces images that are excessively smooth and biased towards the initial estimate of the algorithm. To limit the noise accumulation we propose the use of the pinhole median root prior (PH-MRP) reconstruction algorithm. MRP is a Bayesian reconstruction method that has already been used in PET imaging and shown to possess good noise reduction and edge preservation properties. In this study the PH-MRP algorithm was accelerated with the ordered subsets (OS) procedure and compared to the FBP, OS-EM and conventional Bayesian reconstruction methods in terms of noise reduction, quantitative accuracy, edge preservation and visual quality. The results showed that the accelerated PH-MRP algorithm was very robust. It provided visually pleasing images with lower noise level than the FBP or OS-EM and with smaller bias and sharper edges than the conventional Bayesian methods.

Algorithms↗

Manhattan world: orientation and outlier detection by Bayesian inference.

This letter argues that many visual scenes are based on a "Manhattan" three-dimensional grid that imposes regularities on the image statistics. We construct a Bayesian model that implements this assumption and estimates the viewer orientation relative to the Manhattan grid. For many images, these estimates are good approximations to the viewer orientation (as estimated manually by the authors). These estimates also make it easy to detect outlier structures that are unaligned to the grid. To determine the applicability of the Manhattan world model, we implement a null hypothesis model that assumes that the image statistics are independent of any three-dimensional scene structure. We then use the log-likelihood ratio test to determine whether an image satisfies the Manhattan world assumption. Our results show that if an image is estimated to be Manhattan, then the Bayesian model's estimates of viewer direction are almost always accurate (according to our manual estimates), and vice versa.

Artificial Intelligence↗

Continuous trees and NEVADA simulation: a quadrature approach to modeling continuous random variables in decision analysis.

This paper introduces an improved technique for modeling risk and decision problems that have continuous random variables and probabilistic dependence. Variables are modeled with mixtures of four-parameter random variables, called "continuous trees." Functions of random variables are calculated using gaussian quadrature in a manner called "Nevada simulation" (NumErical Integration of Variance And probabilistic Dependence Analyzer). This technique is compared with traditional decision-tree modeling in terms of analytic technique, solution-time complexity, and accuracy. Nevada simulation takes advantage of the probabilistic independence in a decision problem while allowing for probabilistic dependence to achieve polynomial computational-time complexity for many decision problems. It improves on the accuracy of traditional decision trees by employing larger approximations than traditional decision analysis. It improves on traditional decision analysis by modeling continuous variables with continuous, rather than discrete, distributions. A Bayesian analysis using a mixed discrete-continuous probability distribution for cigarette smoking rate is presented.

Algorithms↗

A physician-based architecture for the construction and use of statistical models.

Physicians need specially tailored computer tools to take advantage of published research results. We present a knowledge-based computer framework--the physician-based (PB) architecture--for constructing such tools, and we use the problem of physicians' interpretation of two-arm parallel randomized clinical trials (TAPRCT) as a working example. Statistical models are represented by influence diagrams. The interpretation of influence-diagram elements are mapped into users' language in a domain-specific, physician-based user interface, called a patient-flow diagram. Statistical-model transformations that maintain the semantic relationships of the model and that embody clinical-epidemiological knowledge are encoded in a mediating structure called the cohort-state diagram. The algorithm that coordinates the interactions among the knowledge representations uses modular actions called construction steps. This architecture has been implemented in a Bayesian system, called THOMAS, that supports physician decision making in light of TAPRCT data. This support entails assessing clinical significance, prior beliefs, and methodological concerns. We suggest that the PB architecture applies to a wide range of statistical tools and users.

Algorithms↗

Joint learning of gene functions--a Bayesian network model approach.

In this paper, we develop a machine learning system for determining gene functions from heterogeneous data sources using a Weighted Naive Bayesian network (WNB). The knowledge of gene functions is crucial for understanding many fundamental biological mechanisms such as regulatory pathways, cell cycles and diseases. Our major goal is to accurately infer functions of putative genes or Open Reading Frames (ORFs) from existing databases using computational methods. However, this task is intrinsically difficult since the underlying biological processes represent complex interactions of multiple entities. Therefore, many functional links would be missing when only one or two sources of data are used in the prediction. Our hypothesis is that integrating evidence from multiple and complementary sources could significantly improve the prediction accuracy. In this paper, our experimental results not only suggest that the above hypothesis is valid, but also provide guidelines for using the WNB system for data collection, training and predictions. The combined training data sets contain information from gene annotations, gene expressions, clustering outputs, keyword annotations, and sequence homology from public databases. The current system is trained and tested on the genes of budding yeast Saccharomyces cerevisiae. Our WNB model can also be used to analyze the contribution of each source of information toward the prediction performance through the weight training process. The contribution analysis could potentially lead to significant scientific discovery by facilitating the interpretation and understanding of the complex relationships between biological entities.

Artificial Intelligence↗

The empirical bias of estimates by restricted maximum likelihood, Bayesian method, and method R under selection for additive, maternal, and dominance models.

Bayesian analysis via Gibbs sampling, restricted maximum likelihood (REML), and Method R were used to estimate variance components for several models of simulated data. Four simulated data sets that included direct genetic effects and different combinations of maternal, permanent environmental, and dominance effects were used. Parents were selected randomly, on phenotype across or within contemporary groups, or on BLUP of genetic value. Estimates by Bayesian analysis and REML were always empirically unbiased in large data sets. Estimates by Method R were biased only with phenotypic selection across contemporary groups; estimates of the additive variance were biased upward, and all the other estimates were biased downward. No empirical bias was observed for Method R under selection within contemporary groups or in data without contemporary group effects. The bias of Method R estimates in small data sets was evaluated using a simple direct additive model. Method R gave biased estimates in small data sets in all types of selection except BLUP. In populations where the selection is based on BLUP of genetic value or where phenotypic selection is practiced mostly within contemporary groups, estimates by Method R are likely to be unbiased. In this case, Method R is an alternative to single-trait REML and Bayesian analysis for analyses of large data sets when the other methods are too expensive to apply.

Animals↗

A simple approach to fitting Bayesian survival models.

There has been much recent work on Bayesian approaches to survival analysis, incorporating features such as flexible baseline hazards, time-dependent covariate effects, and random effects. Some of the proposed methods are quite complicated to implement, and we argue that as good or better results can be obtained via simpler methods. In particular, the normal approximation to the log-gamma distribution yields easy and efficient computational methods in the face of simple multivariate normal priors for baseline log-hazards and time-dependent covariate effects. While the basic method applies to piecewise-constant hazards and covariate effects, it is easy to apply importance sampling to consider smoother functions.

Bayes Theorem↗

A decision-driven system to collect the patient history.

We have developed a computer-administered history designed to directly interview hospitalized patients with pulmonary disease. A frame-based decision system is used to direct the history and to generate a one- to five-member differential diagnostic list based on this history. This system incorporates a cognitive model of question selection and a Bayesian scoring algorithm. Structures to control the choice of questions are embedded in the diagnostic frames and in a QUERY program that makes the final choice of questions. We have compared the behavior of this decision-driven approach with a history taken using a paper questionnaire. The paper-based history presents 182 questions to every patient and captured 75% of 85 pulmonary diseases in its differential lists. The decision-driven system asks 50.7 +/- 31.0 (mean +/- standard deviation) and captured 74% of 61 pulmonary diseases. Our experience suggests that the use of a computerized diagnostic knowledge base to direct the selection of pertinent questions can substantially reduce the number of questions necessary to collect a diagnostically useful patient history.

Artificial Intelligence↗

Prior specification in Bayesian statistics: three cautionary tales.

One of the most important differences between Bayesian and traditional techniques is that the former combines information available beforehand-captured in the prior distribution and reflecting the subjective state of belief before an experiment is carried out-and what the data teach us, as expressed in the likelihood function. Bayesian inference is based on the combination of prior and current information which is reflected in the posterior distribution. The fast growing implementation of Bayesian analysis techniques can be attributed to the development of fast computers and the availability of easy to use software. It has long been established that the specification of prior distributions should receive a lot of attention. Unfortunately, flat distributions are often (inappropriately) used in an automatic fashion in a wide range of types of models. We reiterate that the specification of the prior distribution should be done with great care and support this through three examples. Even in the absence of strong prior information, prior specification should be done at the appropriate scale of biological interest. This often requires incorporation of (weak) prior information based on common biological sense. Very weak and uninformative priors at one scale of the model may result in relatively strong priors at other levels affecting the posterior distribution. We present three different examples intuïvely illustrating this phenomenon indicating that this bias can be substantial (especially in small samples) and is widely present. We argue that complete ignorance or absence of prior information may not exist. Because the central theme of the Bayesian paradigm is to combine prior information with current data, authors should be encouraged to publish their raw data such that every scientist is able to perform an analysis incorporating his/her own (subjective) prior distributions.

Animals↗

The design and construction of a medical simulation model.

This paper describes the design, construction and validation of a probabilistic simulation model of patients who present with abdominal pain. The model incorporates text-book medical knowledge, clinical judgment, and statistics collected from real cases. The knowledge representation combines techniques of Bayesian network modelling with ideas of logistic discrimination. The model is shown to generate convincing, realistic cases; large numbers of artificial cases with no missing observations can be generated quickly. This should make the model a useful tool for investigating factors which limit achievable computer accuracy in the diagnosis of abdominal pain.

Abdominal Pain↗

An empirical Bayesian significance test of cDNA library data.

Automated high-throughput sequencing of cDNA clones from numerous libraries has generated a wealth of information about both genome sequence and relative transcript abundances. A common statistical challenge in the analysis of library sequences is to infer whether there is differential expression for the same transcript under two different conditions, such as normal and diseased tissue. In contrast to the continuously variable intensity measurements from microarray experiments, data from cDNA library sequencing presents itself as a discrete count of the incidence of some clone or transcript in a finite sample. In this paper, we first propose a statistical model for data generated from cDNA library sequencing efforts. The model is based on the Poisson mixed with generalized inverse Gaussian (PGIG), introduced by Sichel (1971, 1975). PGIG has been used in modeling population abundance, ecological studies, word frequencies in publications, etc. Using data from the literature, we show that the proposed model provides a good fit to the observed data. Using this new model for cDNA library data, we developed an empirical Bayesian significance test (EBST) for inferring the statistical significance of differential gene expression from discrete data.

Bayes Theorem↗

Designing for nonparametric Bayesian survival analysis using historical controls.

This paper gives a method for choosing the number of patients N0 out of N available patients to be randomized to current controls om a two-arm study when comparison of nonparametric survival curves is the anticipated method of data analysis. The criterion imposed is that of choosing N0 to minimize the posterior variance of the difference between the current control and experimental survival curves. A nonparametric Bayesian argument incorporating the survival curve of available historical controls establishes the criterion. Formulas and tables which facilitate this computation are presented.

Bayes Theorem↗