Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Informative structure priors: joint learning of dynamic regulatory networks from multiple types of data.

We present a method for jointly learning dynamic models of transcriptional regulatory networks from gene expression data and transcription factor binding location data. Models are automatically learned using dynamic Bayesian network inference algorithms; joint learning is accomplished by incorporating evidence from gene expression data through the likelihood, and from transcription factor binding location data through the prior. We propose a new informative structure prior with two advantages. First, the prior incorporates evidence from location data probabilistically, allowing it to be weighed against evidence from expression data. Second, the prior takes on a factorable form that is computationally efficient when learning dynamic regulatory networks. Results obtained from both simulated and experimental data from the yeast cell cycle demonstrate that this joint learning algorithm can recover dynamic regulatory networks from multiple types of data that are more accurate than those recovered from each type of data in isolation.

Bayes Theorem↗

Predicting analysis times in randomized clinical trials.

Randomized clinical trial designs commonly include one or more planned interim analyses. At these times an external monitoring committee reviews the accumulated data and determines whether it is scientifically and ethically appropriate for the study to continue. With failure-time endpoints, it is common to schedule analyses at the times of occurrence of specified landmark events, such as the 50th event, the 100th event, and so on. Because interim analyses can impose considerable logistical burdens, it is worthwhile predicting their timing as accurately as possible. We describe two model-based methods for making such predictions during the course of a trial. First, we obtain a point prediction by extrapolating the cumulative mortality into the future and selecting the date when the expected number of deaths is equal to the landmark number. Second, we use a Bayesian simulation scheme to generate a predictive distribution of milestone times; prediction intervals are quantiles of this distribution. We illustrate our method with an analysis of data from a trial of immunotherapy in the treatment of chronic granulomatous disease.

Bayes Theorem↗

Model-based detection of lung nodules in computed tomography exams. Thoracic computer-aided diagnosis.

RATIONALE AND OBJECTIVES: In this study, we developed a prototype model-based computer aided detection (CAD) system designed to automatically detect both solid and subsolid pulmonary nodules in computed tomography (CT) images. By using this CAD algorithm, along with the radiologist's initial interpretation, we aim to improve the sensitivity of radiologic readings of CT lung exams. MATERIALS AND METHODS: We have developed a model-based CAD algorithm through the use of precise mathematic models that capture scanner physics and anatomic information. Our model-based CAD algorithm uses multiple segmentation algorithms to extract noteworthy structures in the lungs and a Bayesian statistical model selection framework to determine the probability of various anatomical events throughout the lung. We tested this algorithm on 50 low-dose CT lung cancer screening cases in which ground truth was produced through readings by three expert chest radiologists. RESULTS: Using this model-based CAD algorithm on 50 low-dose CT cases, we measured potential sensitivity improvements of 7% and 5% in two radiologists with respect to all noncalcified nodules, solid and subsolid, greater than 5 mm in diameter. The third radiologist did not miss any nodules in the ground truth set. The CAD algorithm produced 8.3 false positives per case. CONCLUSION: Our prototype CAD system demonstrates promising results as a tool to improve the quality of radiologic readings by increasing radiologist sensitivity. A significant advantage of this model-based approach is that it can be easily extended to support additional anatomic models as clinical understanding and scanning practices improve.

Algorithms↗

Wavelet-based statistical approach for speckle reduction in medical ultrasound images.

A novel speckle-reduction method is introduced, based on soft thresholding of the wavelet coefficients of a logarithmically transformed medical ultrasound image. The method is based on the generalised Gaussian distributed (GGD) modelling of sub-band coefficients. The method used was a variant of the recently published BayesShrink method by Chang and Vetterli, derived in the Bayesian framework for denoising natural images. It was scale adaptive, because the parameters required for estimating the threshold depend on scale and sub-band data. The threshold was computed by Ksigma2/sigma(x), where sigma and sigma(x) were the standard deviation of the noise and the sub-band data of the noise-free image, respectively, and K was a scale parameter. Experimental results showed that the proposed method outperformed the median filter and the homomorphic Wiener filter by 29% in terms of the coefficient of correlation and 4% in terms of the edge preservation parameter. The numerical values of these quantitative parameters indicated the good feature preservation performance of the algorithm, as desired for better diagnosis in medical image processing.

Algorithms↗

Fitting genetic models using Markov Chain Monte Carlo algorithms with BUGS.

Maximum likelihood estimation techniques are widely used in twin and family studies, but soon reach computational boundaries when applied to highly complex models (e.g., models including gene-by-environment interaction and gene-environment correlation, item response theory measurement models, repeated measures, longitudinal structures, extended pedigrees). Markov Chain Monte Carlo (MCMC) algorithms are very well suited to fit complex models with hierarchically structured data. This article introduces the key concepts of Bayesian inference and MCMC parameter estimation and provides a number of scripts describing relatively simple models to be estimated by the freely obtainable BUGS software. In addition, inference using BUGS is illustrated using a data set on follicle-stimulating hormone and luteinizing hormone levels with repeated measures. The examples provided can serve as stepping stones for more complicated models, tailored to the specific needs of the individual researcher.

Algorithms↗

Spatio-temporal autoregressive models defined over brain manifolds.

Multivariate Autoregressive time series models (MAR) are an increasingly used tool for exploring functional connectivity in Neuroimaging. They provide the framework for analyzing the Granger Causality of a given brain region on others. In this article, we shall limit our attention to linear MAR models, in which a set of matrices of autoregressive coefficients Ak (k = 1,...,p) describe the dependence of present values of the image on lagged values of its past. Methods for estimating the Ak and determining which elements that are zero are well-known and are the basis for directed measures of influence. However, to date, MAR models are limited in the number of time series they can handle, forcing the a priori selection of a (small) number of voxels or regions of interest for analysis. This ignores the full spatio-temporal nature of functional brain data which are, in fact, collections of time series sampled over an underlying continuous spatial manifold the brain. A fully spatio-temporal MAR model (ST-MAR) is developed within the framework of functional data analysis. For spatial data, each row of a matrix Ak is the influence field of a given voxel. A Bayesian ST-MAR model is specified in which the influence fields for all voxels are required to vary smoothly over space. This requirement is enforced by penalizing the spatial roughness of the influence fields. This roughness is calculated with a discrete version of the spatial Laplacian operator. A massive reduction in dimensionality of computations is achieved via the singular value decomposition, making an interactive exploration of the model feasible. Use of the model is illustrated with an fMRI time series that was gathered concurrently with EEG in order to analyze the origin of resting brain rhythms.

Bayes Theorem↗

Data-source effects on the sensitivities and specificities of clinical features in the diagnosis of rheumatoid arthritis: the relevance of multiple sources of knowledge for a decision-support system.

An experimental computer system was developed to support diagnosis of rheumatic disorders by computing diagnostic probabilities using modified likelihood ratios. The authors examined whether the performance of the model was affected by the settings in which the data used to derive the likelihood ratios were collected. The sensitivities and specificities of various clinical features for diagnosing rheumatoid arthritis (RA) were obtained from: 1) a study of 1,570 consecutive outpatients at a rheumatology clinic; 2) a review of the literature; 3) estimates by rheumatologists; and 4) a population study. Considerable variations in sensitivity and specificity but satisfactory agreement in likelihood ratios were found across the four data sets. The likelihood ratios were then used to compute the probabilities of RA in a test series of 570 of the rheumatology clinic outpatients. The model's diagnoses with likelihood ratios from the other sources were adequate. When the likelihood ratios from these sources were combined, discrimination came close to what could be achieved by using the likelihood ratios based on the data from the clinic. The method applied in the study, which makes use of variation of input data instead of variation of test series, and the results are relevant to assessing the external validity and transferability of Bayesian decision-support systems.

Adolescent↗

Classical and Bayesian inference in neuroimaging: applications.

In Friston et al. ((2002) Neuroimage 16: 465-483) we introduced empirical Bayes as a potentially useful way to estimate and make inferences about effects in hierarchical models. In this paper we present a series of models that exemplify the diversity of problems that can be addressed within this framework. In hierarchical linear observation models, both classical and empirical Bayesian approaches can be framed in terms of covariance component estimation (e.g., variance partitioning). To illustrate the use of the expectation-maximization (EM) algorithm in covariance component estimation we focus first on two important problems in fMRI: nonsphericity induced by (i) serial or temporal correlations among errors and (ii) variance components caused by the hierarchical nature of multisubject studies. In hierarchical observation models, variance components at higher levels can be used as constraints on the parameter estimates of lower levels. This enables the use of parametric empirical Bayesian (PEB) estimators, as distinct from classical maximum likelihood (ML) estimates. We develop this distinction to address: (i) The difference between response estimates based on ML and the conditional means from a Bayesian approach and the implications for estimates of intersubject variability. (ii) The relationship between fixed- and random-effect analyses. (iii) The specificity and sensitivity of Bayesian inference and, finally, (iv) the relative importance of the number of scans and subjects. The forgoing is concerned with within- and between-subject variability in multisubject hierarchical fMRI studies. In the second half of this paper we turn to Bayesian inference at the first (within-voxel) level, using PET data to show how priors can be derived from the (between-voxel) distribution of activations over the brain. This application uses exactly the same ideas and formalism but, in this instance, the second level is provided by observations over voxels as opposed to subjects. The ensuing posterior probability maps (PPMs) have enhanced anatomical precision and greater face validity, in relation to underlying anatomy. Furthermore, in comparison to conventional SPMs they are not confounded by the multiple comparison problem that, in a classical context, dictates high thresholds and low sensitivity. We conclude with some general comments on Bayesian approaches to image analysis and on some unresolved issues.

Algorithms↗

Molecular systematics of the Jacks (Perciformes: Carangidae) based on mitochondrial cytochrome b sequences using parsimony, likelihood, and Bayesian approaches.

The Carangidae represent a diverse family of marine fishes that include both ecologically and economically important species. Currently, there are four recognized tribes within the family, but phylogenetic relationships among them based on morphology are not resolved. In addition, the tribe Carangini contains species with a variety of body forms and no study has tried to interpret the evolution of this diversity. We used DNA sequences from the mitochondrial cytochrome b gene to reconstruct the phylogenetic history of 50 species from each of the four tribes of Carangidae and four carangoid outgroup taxa. We found support for the monophyly of three tribes within the Carangidae (Carangini, Naucratini, and Trachinotini); however, monophyly of the fourth tribe (Scomberoidini) remains questionable. A sister group relationship between the Carangini and the Naucratini is well supported. This clade is apparently sister to the Trachinotini plus Scomberoidini but there is uncertain support for this relationship. Additionally, we examined the evolution of body form within the tribe Carangini and determined that each of the predominant clades has a distinct evolutionary trend in body form. We tested three methods of phylogenetic inference, parsimony, maximum-likelihood, and Bayesian inference. Whereas the three analyses produced largely congruent hypotheses, they differed in several important relationships. Maximum-likelihood and Bayesian methods produced hypotheses with higher support values for deep branches. The Bayesian analysis was computationally much faster and yet produced phylogenetic hypotheses that were very similar to those of the maximum-likelihood analysis.

Animals↗

Multiple-objective designs in a dose-response experiment.

This article is an extension of the work of Huang and Wong (1), who found dual-objective designs for models with a continuous outcome. We consider quantal dose-response experiments with a binary outcome and develop multiple-objective designs for two or more Bayesian optimality criteria. Using the logit model as an illustrative example, we construct numerically optimal designs for estimating model parameters and percentiles, with possibly unequal interest in each of the objectives. We also show that the popular equal dosage assignment rule can be a rather inefficient design for estimating model parameters under a Bayesian setup.

Bayes Theorem↗

Handling non-negativity in deconvolution of physiological signals: a nonlinear stochastic approach.

A stochastic interpretation of Tikhonov regularization has been recently proposed to attack some open problems of deconvolution when dealing with physiological systems, i.e., in addition to ill-conditioning, infrequent and nonuniform sampling and necessity of having credible confidence intervals. However, the possible violation of the non-negativity constraint cannot be dealt with on firm statistical grounds, since the model of the unknown signal is compatible with negative realizations. In this paper, we propose a new model of the unknown input which excludes negative values. The model is embedded within a Bayesian estimation framework to calculate, by resorting to a Markov chain Monte Carlo algorithm, a nonlinear estimate of the unknown input given by its a posteriori expected value. Applications to simulated and real hormone secretion/pharmacokinetic problems are presented which show that this nonlinear approach is more accurate than the linear one. In addition, more realistic confidence intervals are obtained.

Algorithms↗

Dynamic conditionally linear mixed models for longitudinal data.

We develop a new class of models, dynamic conditionally linear mixed models, for longitudinal data by decomposing the within-subject covariance matrix using a special Cholesky decomposition. Here 'dynamic' means using past responses as covariates and 'conditional linearity' means that parameters entering the model linearly may be random, but nonlinear parameters are nonrandom. This setup offers several advantages and is surprisingly similar to models obtained from the first-order linearization method applied to nonlinear mixed models. First, it allows for flexible and computationally tractable models that include a wide array of covariance structures; these structures may depend on covariates and hence may differ across subjects. This class of models includes, e.g., all standard linear mixed models, antedependence models, and Vonesh-Carter models. Second, it guarantees the fitted marginal covariance matrix of the data is positive definite. We develop methods for Bayesian inference and motivate the usefulness of these models using a series of longitudinal depression studies for which the features of these new models are well suited.

Antidepressive Agents↗

Inferring gene transcriptional modulatory relations: a genetical genomics approach.

Bayesian network modeling is a promising approach to define and evaluate gene expression circuits in diverse tissues and cell types under different experimental conditions. The power and practicality of this approach can be improved by restricting the number of potential interactions among genes and by defining causal relations before evaluating posterior probabilities for billions of networks. A newly developed genetical genomics method that combines transcriptome profiling with complex trait analysis now provides strong constraints on network architecture. This method detects those chromosomal intervals responsible for differences in mRNA expression using quantitative trait locus (QTL) mapping. We have developed an efficient Bayesian approach that exploits the genetical genomics method to focus computational effort on the most plausible gene modulatory networks. We exploit a dense marker map for a genetic reference population (GRP) that consists of 32 BXD strains of mice made by intercrossing two progenitor strains--C57BL/6J and DBA/2J. These progenitors differ at approximately 1.3 million known single nucleotide polymorphisms (SNPs), all of which can be exploited to estimate the probability that a gene contains functional polymorphisms that segregate within the GRP. We constructed 66 candidate networks that include all the candidate modulator genes located in the 209 statistically significant trans-acting QTL regions. SNPs that distinguish between the two progenitor strains were used to further winnow the list of candidate modulators. Bayesian network was then used to identify the genetic modulatory relations that best explain the microarray data.

Algorithms↗

CMfinder--a covariance model based RNA motif finding algorithm.

MOTIVATION: The recent discoveries of large numbers of non-coding RNAs and computational advances in genome-scale RNA search create a need for tools for automatic, high quality identification and characterization of conserved RNA motifs that can be readily used for database search. Previous tools fall short of this goal. RESULTS: CMfinder is a new tool to predict RNA motifs in unaligned sequences. It is an expectation maximization algorithm using covariance models for motif description, featuring novel integration of multiple techniques for effective search of motif space, and a Bayesian framework that blends mutual information-based and folding energy-based approaches to predict structure in a principled way. Extensive tests show that our method works well on datasets with either low or high sequence similarity, is robust to inclusion of lengthy extraneous flanking sequence and/or completely unrelated sequences, and is reasonably fast and scalable. In testing on 19 known ncRNA families, including some difficult cases with poor sequence conservation and large indels, our method demonstrates excellent average per-base-pair accuracy--79% compared with at most 60% for alternative methods. More importantly, the resulting probabilistic model can be directly used for homology search, allowing iterative refinement of structural models based on additional homologs. We have used this approach to obtain highly accurate covariance models of known RNA motifs based on small numbers of related sequences, which identified homologs in deeply-diverged species.

Algorithms↗

Bayesian restoration of a hidden Markov chain with applications to DNA sequencing.

Hidden Markov models (HMMs) are a class of stochastic models that have proven to be powerful tools for the analysis of molecular sequence data. A hidden Markov model can be viewed as a black box that generates sequences of observations. The unobservable internal state of the box is stochastic and is determined by a finite state Markov chain. The observable output is stochastic with distribution determined by the state of the hidden Markov chain. We present a Bayesian solution to the problem of restoring the sequence of states visited by the hidden Markov chain from a given sequence of observed outputs. Our approach is based on a Monte Carlo Markov chain algorithm that allows us to draw samples from the full posterior distribution of the hidden Markov chain paths. The problem of estimating the probability of individual paths and the associated Monte Carlo error of these estimates is addressed. The method is illustrated by considering a problem of DNA sequence multiple alignment. The special structure for the hidden Markov model used in the sequence alignment problem is considered in detail. In conclusion, we discuss certain interesting aspects of biological sequence alignments that become accessible through the Bayesian approach to HMM restoration.

Algorithms↗

Maximizing sensitivity in medical diagnosis using biased minimax probability machine.

The challenging task of medical diagnosis based on machine learning techniques requires an inherent bias, i.e., the diagnosis should favor the "ill" class over the "healthy" class, since misdiagnosing a patient as a healthy person may delay the therapy and aggravate the illness. Therefore, the objective in this task is not to improve the overall accuracy of the classification, but to focus on improving the sensitivity (the accuracy of the "ill" class) while maintaining an acceptable specificity (the accuracy of the "healthy" class). Some current methods adopt roundabout ways to impose a certain bias toward the important class, i.e., they try to utilize some intermediate factors to influence the classification. However, it remains uncertain whether these methods can improve the classification performance systematically. In this paper, by engaging a novel learning tool, the biased minimax probability machine (BMPM), we deal with the issue in a more elegant way and directly achieve the objective of appropriate medical diagnosis. More specifically, the BMPM directly controls the worst case accuracies to incorporate a bias toward the "ill" class. Moreover, in a distribution-free way, the BMPM derives the decision rule in such a way as to maximize the worst case sensitivity while maintaining an acceptable worst case specificity. By directly controlling the accuracies, the BMPM provides a more rigorous way to handle medical diagnosis; by deriving a distribution-free decision rule, the BMPM distinguishes itself from a large family of classifiers, namely, the generative classifiers, where an assumption on the data distribution is necessary. We evaluate the performance of the model and compare it with three traditional classifiers: the k-nearest neighbor, the naive Bayesian, and the C4.5. The test results on two medical datasets, the breast-cancer dataset and the heart disease dataset, show that the BMPM outperforms the other three models.

Algorithms↗

SIMMAP: stochastic character mapping of discrete traits on phylogenies.

BACKGROUND: Character mapping on phylogenies has played an important, if not critical role, in our understanding of molecular, morphological, and behavioral evolution. Until very recently we have relied on parsimony to infer character changes. Parsimony has a number of serious limitations that are drawbacks to our understanding. Recent statistical methods have been developed that free us from these limitations enabling us to overcome the problems of parsimony by accommodating uncertainty in evolutionary time, ancestral states, and the phylogeny. RESULTS: SIMMAP has been developed to implement stochastic character mapping that is useful to both molecular evolutionists, systematists, and bioinformaticians. Researchers can address questions about positive selection, patterns of amino acid substitution, character association, and patterns of morphological evolution. CONCLUSION: Stochastic character mapping, as implemented in the SIMMAP software, enables users to address questions that require mapping characters onto phylogenies using a probabilistic approach that does not rely on parsimony. Analyses can be performed using a fully Bayesian approach that is not reliant on considering a single topology, set of substitution model parameters, or reconstruction of ancestral states. Uncertainty in these quantities is accommodated by using MCMC samples from their respective posterior distributions.

Animals↗

Bayesian estimation of range for microsatellite loci.

Microsatellite loci have become important in population genetics because of their high level of polymorphism in natural populations, very frequent occurrence throughout the genome, and apparently high mutation rate. Observed repeat numbers (alleles size) in natural populations and expectations based on computer simulations suggest that the range of repeat numbers at a microsatellite locus is restricted. This range is a key parameter that should be properly estimated in order to proceed with calculations of divergence times in phylogenetic studies and to better investigate the within- and between-population variability. The 'plug-in' estimate of range based on the minimum and maximum value observed in a sample is not satisfactory because of the relatively large number of alleles in comparison with typical sample sizes. In this paper, a set of data from 30 dinucleotide microsatellite loci is analysed under the assumption of independence among loci. Bayesian inference on range for one locus is obtained by assuming that constraints on range values exist as sharp bounds. Closed-form calculations and robustness revealed by our analysis suggest that the proposed Bayesian approach might be routinely used by researchers to classify microsatellite loci according to the estimated value of their allelic range.

Bayes Theorem↗