Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

Fiber tracking from DTI using linear state space models: detectability of the pyramidal tract.

Diffusion tensor imaging (DTI) is an emerging and promising tool to provide information about the course of white matter fiber tracts in the human brain. Based on specific acquisition schemes, diffusion tensor data resemble local fiber orientations allowing for a reconstruction of the fiber bundles. Current techniques to calculate fascicles range from simple heuristic tracking solutions to Bayesian and differential equations approaches. Most methods are based only on local diffusion information, often resulting in bending or kinking fiber paths in voxels with reduced diffusion properties. In this article we present a new tracking approach based on linear state space models encompassing an inherent smoothness criterion to avoid too wiggly tracked fiber bundles. The new technique will be described formally and tested on simulated and real data. The performance tests are focused on the pyramidal tract, where we employed a test-retest study and a group comparison in healthy subjects. Anatomical course was confirmed in a patient with selective degeneration of the pyramidal tract. The potential of the presented technique for improved neurosurgical planning is demonstrated by visualization of a tumor-induced displacement of the motor pathways. The paper closes with a thorough discussion of perspectives and limitations of the new tracking approach.

Adult↗

Bayesian analysis, pattern analysis, and data mining in health care.

PURPOSE OF REVIEW: To discuss the current role of data mining and Bayesian methods in biomedicine and heath care, in particular critical care. RECENT FINDINGS: Bayesian networks and other probabilistic graphical models are beginning to emerge as methods for discovering patterns in biomedical data and also as a basis for the representation of the uncertainties underlying clinical decision-making. At the same time, techniques from machine learning are being used to solve biomedical and health-care problems. SUMMARY: With the increasing availability of biomedical and health-care data with a wide range of characteristics there is an increasing need to use methods which allow modeling the uncertainties that come with the problem, are capable of dealing with missing data, allow integrating data from various sources, explicitly indicate statistical dependence and independence, and allow integrating biomedical and clinical background knowledge. These requirements have given rise to an influx of new methods into the field of data analysis in health care, in particular from the fields of machine learning and probabilistic graphical models.

Bayes Theorem↗

Reassessing benzene risks using internal doses and Monte-Carlo uncertainty analysis.

Human cancer risks from benzene have been estimated from epidemiological data, with supporting evidence from animal bioassay data. This article reexamines the animal-based risk assessments using physiologically based pharmacokinetic (PBPK) models of benzene metabolism in animals and humans. Internal doses (total benzene metabolites) from oral gavage experiments in mice are well predicted by the PBPK model. Both the data and the PBPK model outputs are also well described by a simple nonlinear (Michaelis-Menten) regression model, as previously used by Bailer and Hoel [Metabolite-based internal doses used in risk assessment of benzene. Environ Health Perspect 82:177-184 (1989)]. Refitting the multistage model family to internal doses changes the maximum-likelihood estimate (MLE) dose-response curve for mice from linear-quadratic to purely cubic, so that low-dose risk estimates are smaller than in previous risk assessments. In contrast to Bailer and Hoel's findings using interspecies dose conversion, the use of internal dose estimates for humans from a PBPK model reduces estimated human risks at low doses. Sensitivity analyses suggest that the finding of a nonlinear MLE dose-response curve at low doses is robust to changes in internal dose definitions and more consistent with epidemiological data than earlier risk models. A Monte-Carlo uncertainty analysis based on maximum-entropy probabilities and Bayesian conditioning is used to develop an entire probability distribution for the true but unknown dose-response function. This allows the probability of a positive low-dose slope to be quantified: It is about 10%. An upper 95% confidence limit on the low-dose slope of excess risk is also obtained directly from the posterior distribution and is similar to previous q1* values. This approach suggests that the excess risk due to benzene exposure may be nonexistent (or even negative) at sufficiently low doses. Two types of biological information about benzene effects--pharmacokinetic and hematotoxic--are examined to test the plausibility of this finding. A framework for incorporating causally relevant biological information into benzene risk assessment is introduced, and it is shown that both pharmacokinetic and hematotoxic models appear to be consistent with the hypothesis that sufficiently low concentrations of inhaled benzene do not create and excess risk.

Administration, Oral↗

Potts model for haplotype associations.

Bayesian spatial modeling has become important in disease mapping and has also been suggested as a useful tool in genetic fine mapping. We have implemented the Potts model and applied it to the Genetic Analysis Workshop 14 (GAW14) simulated data. Because the "answers" were known we have analyzed latent phenotype P1-related observed phenotypes affection status (genetically determined) and i (random) in the Danacaa population replicate 2. Analysis of the microsatellite/single-nucleotide polymorphism-based haplotypes at chromosomes 1 and 3 failed to identify multiple clusters of haplotype effects. However, the analysis of separately simulated data with postulated differences in the effects of the two clusters has yielded clear estimated division into the two clusters, demonstrating the correctness of the algorithm. Although we could not clearly identify the disease-related and the non-associated groups of haplotypes, results of both GAW14 and our own simulation encourage us to improve the efficiency and sensitivity of the estimation algorithm and to further compare the proposed method with more traditional methods.

Computer Simulation↗

Data mining and computationally intensive methods: summary of Group 7 contributions to Genetic Analysis Workshop 13.

The Framingham Heart Study data, as well as a related simulated data set, were generously provided to the participants of the Genetic Analysis Workshop 13 in order that newly developed and emerging statistical methodologies could be tested on that well-characterized data set. The impetus driving the development of novel methods is to elucidate the contributions of genes, environment, and interactions between and among them, as well as to allow comparison between and validation of methods. The seven papers that comprise this group used data-mining methodologies (tree-based methods, neural networks, discriminant analysis, and Bayesian variable selection) in an attempt to identify the underlying genetics of cardiovascular disease and related traits in the presence of environmental and genetic covariates. Data-mining strategies are gaining popularity because they are extremely flexible and may have greater efficiency and potential in identifying the factors involved in complex disorders. While the methods grouped together here constitute a diverse collection, some papers asked similar questions with very different methods, while others used the same underlying methodology to ask very different questions. This paper briefly describes the data-mining methodologies applied to the Genetic Analysis Workshop 13 data sets and the results of those investigations.

Bayes Theorem↗

Comparison of REML and Gibbs sampling estimates of multi-trait genetic parameters in Scots pine.

Multi-trait (co)variance estimation is an important topic in plant and animal breeding. In this study we compare estimates obtained with restricted maximum likelihood (REML) and Bayesian Gibbs sampling of simulated data and of three traits (diameter, height and branch angle) from a 26-year-old partial diallel progeny test of Scots pine (Pinus sylvestris L.). Based on the results from the simulated data we can conclude that the REML estimates are accurate but the mode of posterior distributions from the Gibbs sampling can be overestimated depending on the level of the heritability. The mean and median of the posteriors were considerably higher than the expected values of the heritabilities. The confidence intervals calculated with the delta method were biased downwardly. The highest probability density (HPD) interval provides a better interval estimate, but could be slightly biased at the lower level. Similar differences between REML and Gibbs sampling estimates were found for the Scots pine data. We conclude that further simulation studies are needed in order to evaluate the effect of different priors on (co)variance components in the genetic individual model.

Analysis of Variance↗

Homomorphic wavelet thresholding technique for denoising medical ultrasound images.

A novel homomorphic wavelet thresholding technique for reducing speckle noise in medical ultrasound images is presented. First, we show that the speckle wavelet coefficients in the logarithmically transformed ultrasound images are best described by the Nakagami family of distributions. By exploiting this speckle model and the Laplacian signal prior, a closed form, data-driven, and spatially adaptive threshold is derived in the Bayesian framework. The spatial adaptivity allows the additional information of the image (such as identification of homogeneous or heterogeneous regions) to be incorporated into the algorithm. Further, the threshold has been extended to the redundant wavelet representation, which yields better results than the decimated wavelet transform. Experimental results demonstrate the improved performance of the proposed method over other well-known speckle reduction filters. The application of the proposed method to a realistic US test image shows that the new technique, named HomoGenThresh, outperforms the best wavelet-based denoising method reported in [1] by more than 1.6 dB, Lee filter by 3.6 dB, Kaun filter by 3.1 dB and band-adaptive soft thresholding [2] by 2.1 dB at an input signal-to-noise ratio (SNR) of 13.6 dB.

Algorithms↗

Intensity-based hierarchical Bayes method improves testing for differentially expressed genes in microarray experiments.

BACKGROUND: The small sample sizes often used for microarray experiments result in poor estimates of variance if each gene is considered independently. Yet accurately estimating variability of gene expression measurements in microarray experiments is essential for correctly identifying differentially expressed genes. Several recently developed methods for testing differential expression of genes utilize hierarchical Bayesian models to "pool" information from multiple genes. We have developed a statistical testing procedure that further improves upon current methods by incorporating the well-documented relationship between the absolute gene expression level and the variance of gene expression measurements into the general empirical Bayes framework. RESULTS: We present a novel Bayesian moderated-T, which we show to perform favorably in simulations, with two real, dual-channel microarray experiments and in two controlled single-channel experiments. In simulations, the new method achieved greater power while correctly estimating the true proportion of false positives, and in the analysis of two publicly-available "spike-in" experiments, the new method performed favorably compared to all tested alternatives. We also applied our method to two experimental datasets and discuss the additional biological insights as revealed by our method in contrast to the others. The R-source code for implementing our algorithm is freely available at http://eh3.uc.edu/ibmt. CONCLUSION: We use a Bayesian hierarchical normal model to define a novel Intensity-Based Moderated T-statistic (IBMT). The method is completely data-dependent using empirical Bayes philosophy to estimate hyperparameters, and thus does not require specification of any free parameters. IBMT has the strength of balancing two important factors in the analysis of microarray data: the degree of independence of variances relative to the degree of identity (i.e. t-tests vs. equal variance assumption), and the relationship between variance and signal intensity. When this variance-intensity relationship is weak or does not exist, IBMT reduces to a previously described moderated t-statistic. Furthermore, our method may be directly applied to any array platform and experimental design. Together, these properties show IBMT to be a valuable option in the analysis of virtually any microarray experiment.

Animals↗

Polygonal and polyhedral contour reconstruction in computed tomography.

This paper is about three-dimensional (3-D) reconstruction of a binary image from its X-ray tomographic data. We study the special case of a compact uniform polyhedron totally included in a uniform background and directly perform the polyhedral surface estimation. We formulate this problem as a nonlinear inverse problem using the Bayesian framework. Vertice estimation is done without using a voxel approximation of the 3-D image. It is based on the construction and optimization of a regularized criterion that accounts for surface smoothness. We investigate original deterministic local algorithms, based on the exact computation of the line projections, their update, and their derivatives with respect to the vertice coordinates. Results are first derived in the two-dimensional (2-D) case, which consists of reconstructing a 2-D object of deformable polygonal contour from its tomographic data. Then, we investigate the 3-D extension that requires technical adaptations. Simulation results illustrate the performance of polygonal and polyhedral reconstruction algorithms in terms of quality and computation time.

Algorithms↗

Locating disease genes using Bayesian variable selection with the Haseman-Elston method.

BACKGROUND: We applied stochastic search variable selection (SSVS), a Bayesian model selection method, to the simulated data of Genetic Analysis Workshop 13. We used SSVS with the revisited Haseman-Elston method to find the markers linked to the loci determining change in cholesterol over time. To study gene-gene interaction (epistasis) and gene-environment interaction, we adopted prior structures, which incorporate the relationship among the predictors. This allows SSVS to search in the model space more efficiently and avoid the less likely models. RESULTS: In applying SSVS, instead of looking at the posterior distribution of each of the candidate models, which is sensitive to the setting of the prior, we ranked the candidate variables (markers) according to their marginal posterior probability, which was shown to be more robust to the prior. Compared with traditional methods that consider one marker at a time, our method considers all markers simultaneously and obtains more favorable results. CONCLUSIONS: We showed that SSVS is a powerful method for identifying linked markers using the Haseman-Elston method, even for weak effects. SSVS is very effective because it does a smart search over the entire model space.

Bayes Theorem↗

A probabilistic Classifier System and its application in data mining.

The article is about a new Classifier System framework for classification tasks called BYP-CS (for BaYesian Predictive Classifier System). The proposed CS approach abandons the focus on high accuracy and addresses a well-posed Data Mining goal, namely, that of uncovering the low-uncertainty patterns of dependence that manifest often in the data. To attain this goal, BYP-CS uses a fair amount of probabilistic machinery, which brings its representation language closer to other related methods of interest in statistics and machine learning. On the practical side, the new algorithm is seen to yield stable learning of compact populations, and these still maintain a respectable amount of predictive power. Furthermore, the emerging rules self-organize in interesting ways, sometimes providing unexpected solutions to certain benchmark problems.

Algorithms↗

Risk-based environmental remediation: Bayesian Monte Carlo analysis and the expected value of sample information.

A methodology that simulates outcomes from future data collection programs, utilizes Bayesian Monte Carlo analysis to predict the resulting reduction in uncertainty in an environmental fate-and-transport model, and estimates the expected value of this reduction in uncertainty to a risk-based environmental remediation decision is illustrated considering polychlorinated biphenyl (PCB) sediment contamination and uptake by winter flounder in New Bedford Harbor, MA. The expected value of sample information (EVSI), the difference between the expected loss of the optimal decision based on the prior uncertainty analysis and the expected loss of the optimal decision from an updated information state, is calculated for several sampling plan. For the illustrative application we have posed, the EVSI for a sampling plan of two data points is $9.4 million, for five data points is $10.4 million, and for ten data points is $11.5 million. The EVSI for sampling plans involving larger numbers of data points is bounded by the expected value of perfect information, $15.6 million. A sensitivity analysis is conducted to examine the effect of selected model structure and parametric assumptions on the optimal decision and the EVSI. The optimal decision (total area to be dredged) is sensitive to the assumption of linearity between PCB sediment concentration and flounder PCB body burden and to the assumed relationship between area dredged and the harbor-wide average sediment PCB concentration; these assumptions also have a moderate impact on the computed EVSI. The EVSI is most sensitive to the unit cost of remediation and rather insensitive to the penalty cost associated with under-remediation.

Animals↗

Markov chain Monte Carlo methods in biostatistics.

Appropriate models in biostatistics are often quite complicated. Such models are typically most easily fit using Bayesian methods, which can often be implemented using simulation techniques. Markov chain Monte Carlo (MCMC) methods are an important set of tools for such simulations. We give an overview and references of this rapidly emerging technology along with a relatively simple example. MCMC techniques can be viewed as extensions of iterative maximization techniques, but with random jumps rather than maximizations at each step. Special care is needed when implementing iterative maximization procedures rather than closed-form methods, and even more care is needed with iterative simulation procedures: it is substantially more difficult to monitor convergence to a distribution than to a point. The most reliable implementations of MCMC build upon results from simpler models fit using combinations of maximization algorithms and noniterative simulations, so that the user has a rough idea of the location and scale of the posterior distribution of the quantities of interest under the more complicated model. These concerns with implementation, however, should not deter the biostatistician from using MCMC methods, but rather help to ensure wise use of these powerful techniques.

Algorithms↗

A Bayesian morphometry algorithm.

Most methods for structure-function analysis of the brain in medical images are usually based on voxel-wise statistical tests performed on registered magnetic resonance (MR) images across subjects. A major drawback of such methods is the inability to accurately locate regions that manifest nonlinear associations with clinical variables. In this paper, we propose Bayesian morphological analysis methods, based on a Bayesian-network representation, for the analysis of MR brain images. First, we describe how Bayesian networks (BNs) can represent probabilistic associations among voxels and clinical (function) variables. Second, we present a model-selection framework, which generates a BN that captures structure-function relationships from MR brain images and function variables. We demonstrate our methods in the context of determining associations between regional brain atrophy (as demonstrated on MR images of the brain), and functional deficits. We employ two data sets for this evaluation: the first contains MR images of 11 subjects, where associations between regional atrophy and a functional deficit are almost linear; the second data set contains MR images of the ventricles of 84 subjects, where the structure-function association is nonlinear. Our methods successfully identify voxel-wise morphological changes that are associated with functional deficits in both data sets, whereas standard statistical analysis (i.e., t-test and paired t-test) fails in the nonlinear-association case.

Aged↗

Shrinkage estimation method for mapping multiple quantitative trait loci.

In this article, shrinkage estimation method for multiple-marker analysis and for mapping multiple quantitative trait loci (QTL) was reviewed. For multiple-marker analysis, Xu (Genetics, 2003, 163:789-801) developed a Bayesian shrinkage estimation (BSE) method. The key to the success of this method is to allow each marker effect have its own variance parameter, which in turn has its own prior distribution so that the variance can be estimated from the data. Under this hierarchical model, a large number of markers can be handled although most of them may have negligible effects. Under epistatic genetic model, however, the running time is very long. To overcome this problem, a novel method of incorporating the idea described above into maximum likelihood, known as penalized likelihood method, was proposed. A simulated study showed that this method can handle a model with multiple effects, which are ten times larger than the sample size. For multiple QTL analysis, two modified versions for the BSE method were introduced: one is the fixed-interval method and another is the variable-interval method. The former deals with markers with intermediate density, and the latter can handle markers with extremely high density as well as model with epistatic effects. For the detection of epistatic effects, penalized likelihood method and the variable-interval approach of the BSE method are available.

Bayes Theorem↗

VAMPIRE microarray suite: a web-based platform for the interpretation of gene expression data.

Microarrays are invaluable high-throughput tools used to snapshot the gene expression profiles of cells and tissues. Among the most basic and fundamental questions asked of microarray data is whether individual genes are significantly activated or repressed by a particular stimulus. We have previously presented two Bayesian statistical methods for this level of analysis, collectively known as variance-modeled posterior inference with regional exponentials (VAMPIRE). These methods each require a sophisticated modeling step followed by integration of a posterior probability density. We present here a publicly available, web-based platform that allows users to easily load data, associate related samples and identify differentially expressed features using the VAMPIRE statistical framework. In addition, this suite of tools seamlessly integrates a novel gene annotation tool, known as GOby, which identifies statistically overrepresented gene groups. Unlike other tools in this genre, GOby can localize enrichment while respecting the hierarchical structure of annotation systems like Gene Ontology (GO). By identifying statistically significant enrichment of GO terms, Kyoto Encyclopedia of Genes and Genomes pathways, and TRANSFAC transcription factor binding sites, users can gain substantial insight into the physiological significance of sets of differentially expressed genes. The VAMPIRE microarray suite can be accessed at http://genome.ucsd.edu/microarray.

Bayes Theorem↗

Identification of patients with impaired hepatic drug metabolism using a limited sampling procedure for estimation of phenazone (antipyrine) pharmacokinetic parameters.

Phenazone (antipyrine) 1g was given by short intravenous infusion to 62 study participants (10 healthy drug-free volunteers and 52 patients with chronic liver disease). A Bayesian approach was developed to determine the individual pharmacokinetic parameters of phenazone. Statistical characteristics of the population pharmacokinetic parameters were first evaluated for 30 patients. When combined with 1 plasma drug concentration from members of the second group, these led to a Bayesian estimation of individual pharmacokinetic parameters for the remaining 32 individuals. Total clearance computed by Bayesian estimation was compared with maximal likelihood estimation of this parameter, the classical procedure. No statistically significant differences were found. Performance of the developed methodology was evaluated by computing bias and precision. The mean error was 0.0477 L/h. The precision of the prediction of this parameter (0.155 L/h) remained lower than the interindividual standard deviation (0.765 L/h). This procedure enables the estimation of individual pharmacokinetic parameters for phenazone. In this study, numerous laboratory tests were performed. A highly significant correlation (p < 0.001) was found between phenazone clearance and the prothrombin time, albumin, gamma-globulin, factor V, antithrombin III, fibrinogen and total bilirubin. Discriminant analysis determined that protein, alkaline phosphatase, creatininaemia and gamma-globulin had more significant discriminating power and gave better prognostic results than those seen with the Child-Pugh test.

Adult↗

Molecular phylogeny of parabasalids inferred from small subunit rRNA sequences, with emphasis on the Hypermastigea.

Small subunit rRNA gene sequences were identified without cultivation from parabasalid symbionts of termites belonging to the hypermastigid orders Trichonymphida (the genera Hoplonympha, Staurojoenina, Teranympha, and Eucomonympha) and Spirotrichonymphida (Spirotrichonymphella), and from four yet-unidentified parabasalid symbionts of the termite Incisitermes minor. All these new sequences were analyzed by Bayesian, likelihood, and parsimony methods in a broad phylogeny including all identified parabasalid sequences available in databases and some as yet unidentified sequences probably derived from hypermastigids. A salient point of our study focused on hypermastigids was the polyphyly of this class. We also noted a clear dichotomy between Trichonymphida and the other parabasalid taxa. However, this hypermastigid order was apparently polyphyletic, probably reflecting its morphological diversity. Among Trichonymphida, Teranympha (Teranymphidae) grouped together with the members of the family Eucomonymphidae, suggesting that its family status is ambiguous. The monophyletic lineage composed by Spirotrichonymphida exhibited a narrower branching pattern than Trichonymphida. The root of parabasalids was examined but could not be discerned accurately.

Animals↗