Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Prediction of contact maps by GIOHMMs and recurrent neural networks using lateral propagation from all four cardinal corners.

MOTIVATION: Accurate prediction of protein contact maps is an important step in computational structural proteomics. Because contact maps provide a translation and rotation invariant topological representation of a protein, they can be used as a fundamental intermediary step in protein structure prediction. RESULTS: We develop a new set of flexible machine learning architectures for the prediction of contact maps, as well as other information processing and pattern recognition tasks. The architectures can be viewed as recurrent neural network implemantations of a class of Bayesian networks we call generalized input-output HMMs (GIOHMMs). For the specific case of contact maps, contextual information is propagated laterally through four hidden planes, one for each cardinal corner. We show that these architectures can be trained from examples and yield contact map predictors that outperform previously reported methods. While several extensions and improvements are in progress, the current version can accurately predict 60.5% of contacts at a distance cutoff of 8 A and 45% of distant contacts at 10 A, for proteins of length up to 300.

Algorithms↗

A two-sample test with interval censored data via multiple imputation.

Interval censored data arise naturally in large scale panel studies where subjects can only be followed periodically and the event of interest can only be recorded as having occurred between two examination times. In this paper we consider the problem of comparing two interval-censored samples. We propose to impute exact failure times from interval-censored observations to obtain right censored data, then apply existing techniques, such as Harrington and Fleming's G(rho) tests to imputed right censored data. To appropriately account for variability, a multiple imputation algorithm based on the approximate Bayesian bootstrap (ABB) is discussed. Through simulation studies we find that it performs well. The advantage of our proposal is its simplicity to implement and adaptability to incorporate many existing two-sample comparison techniques for right censored data. The method is illustrated by reanalysing the Breast Cosmesis Study data set.

Algorithms↗

Lesion size quantification in SPECT using an artificial neural network classification approach.

An artificial neural network (ANN) has been developed to determine the size of lesions detected in single photon emission computed tomographic images. The network is the Learning Vector Quantizer and is trained to perform size quantification based on image neighborhoods extracted around the lesions. The ANN is compared to the optimal, Bayesian algorithm developed to perform the same task using the unreconstructed, projection data. The performance of the neural network is evaluated at two different noise levels. The Bayesian algorithm provides the upper bound for size quantification performance against which the ANN is compared. In the ideal case where the Bayesian algorithm has explicit knowledge of the underlying distributions, its performance is superior to that of the neural network. However, in the more realistic case where the distributions need to be estimated from the same learning sample the ANN was trained on, the two algorithms have comparable performances.

Algorithms↗

Bayesian sparse hidden components analysis for transcription regulation networks.

MOTIVATION: In systems like Escherichia Coli, the abundance of sequence information, gene expression array studies and small scale experiments allows one to reconstruct the regulatory network and to quantify the effects of transcription factors on gene expression. However, this goal can only be achieved if all information sources are used in concert. RESULTS: Our method integrates literature information, DNA sequences and expression arrays. A set of relevant transcription factors is defined on the basis of literature. Sequence data are used to identify potential target genes and the results are used to define a prior distribution on the topology of the regulatory network. A Bayesian hidden component model for the expression array data allows us to identify which of the potential binding sites are actually used by the regulatory proteins in the studied cell conditions, the strength of their control, and their activation profile in a series of experiments. We apply our methodology to 35 expression studies in E.Coli with convincing results. AVAILABILITY: www.genetics.ucla.edu/labs/sabatti/software.html SUPPLEMENTARY INFORMATION: The supplementary material are available at Bioinformatics online.

Algorithms↗

The application of Bayesian techniques in the interpretation of bioassay data.

The inverse problem of internal dosimetry is naturally posed as a problem of Bayesian inference. The Bayesian approach is of practical importance in three areas: (1) avoiding false positives in the detection of rare events, (2) the calculation of uncertainties, and (3) the calculation of multiple intakes, all of which are important for internal dosimetry. In this paper, the Bayesian approach to the interpretation of measurements is first reviewed using a simple conceptual example. Then, a simple 239Pu case using IMBA expert is discussed, and finally a current cutting-edge example is discussed involving real 238Pu data calculated with a Markov Chain Monte Carlo algorithm and with exact calculation of poisson likelihood functions.

Algorithms↗

Model-based region-of-interest selection in dynamic breast MRI.

Magnetic resonance imaging (MRI) is emerging as a powerful tool for the diagnosis of breast abnormalities. Dynamic analysis of the temporal pattern of contrast uptake has been applied in differential diagnosis of benign and malignant lesions to improve specificity. Selecting a region of interest (ROI) is an almost universal step in the process of examining the contrast uptake characteristics of a breast lesion. We propose an ROI selection method that combines model-based clustering of the pixels with Bayesian morphology, a new statistical image segmentation method. We then investigate tools for subsequent analysis of signal intensity time course data in the selected region. Results on a database of 19 patients indicate that the method provides informative segmentations and good detection rates.

Bayes Theorem↗

Construction of an abdominal probabilistic atlas and its application in segmentation.

There have been significant efforts to build a probabilistic atlas of the brain and to use it for many common applications, such as segmentation and registration. Though the work related to brain atlases can be applied to nonbrain organs, less attention has been paid to actually building an atlas for organs other than the brain. Motivated by the automatic identification of normal organs for applications in radiation therapy treatment planning, we present a method to construct a probabilistic atlas of an abdomen consisting of four organs (i.e., liver, kidneys, and spinal cord). Using 32 noncontrast abdominal computed tomography (CT) scans, 31 were mapped onto one individual scan using thin plate spline as the warping transform and mutual information (MI) as the similarity measure. Except for an initial coarse placement of four control points by the operators, the MI-based registration was automatic. Additionally, the four organs in each of the 32 CT data sets were manually segmented. The manual segmentations were warped onto the "standard" patient space using the same transform computed from their gray scale CT data set and a probabilistic atlas was calculated. Then, the atlas was used to aid the segmentation of low-contrast organs in an additional 20 CT data sets not included in the atlas. By incorporating the atlas information into the Bayesian framework, segmentation results clearly showed improvements over a standard unsupervised segmentation method.

Algorithms↗

A survey of current Bayesian gene mapping methods.

Recently, there has been much interest in the use of Bayesian statistical methods for performing genetic analyses. Many of the computational difficulties previously associated with Bayesian analysis, such as multidimensional integration, can now be easily overcome using modern high-speed computers and Markov chain Monte Carlo (MCMC) methods. Much of this new technology has been used to perform gene mapping, especially through the use of multi-locus linkage disequilibrium techniques. This review attempts to summarise some of the currently available methods and the software available to implement these methods.

Bayes Theorem↗

A Bayesian approach to multipoint mapping in nuclear families.

We describe the application of a Markov Chain Monte Carlo approach for multipoint mapping of a quantitative trait locus to the Nuclear Families simulated data. The method involves repeated sampling of genotype vectors for each nuclear family from their conditional distributions, given phenotypes, markers, and model parameters, using peeling and gene dropping, followed by random sampling of each model parameter given genotypes and the other parameters. Reversible jump methods are used to sample the number of trait loci.

Bayes Theorem↗

Mapping multiple quantitative trait Loci for ordinal traits.

Many complex traits in humans and other organisms show ordinal phenotypic variation but do not follow a simple Mendelian pattern of inheritance. These ordinal traits are presumably determined by many factors, including genetic and environmental components. Several statistical approaches to mapping quantitative trait loci (QTL) for such traits have been developed based on a single-QTL model. However, statistical methods for mapping multiple QTL are not well studied as continuous traits. In this paper, we propose a Bayesian method implemented via the Markov chain Monte Carlo (MCMC) algorithm to map multiple QTL for ordinal traits in experimental crosses. We model the ordinal traits under the multiple threshold model, which assumes a latent continuous variable underlying the ordinal phenotypes. The ordinal phenotype and the latent continuous variable are linked through some fixed but unknown thresholds. We adopt a standardized threshold model, which has several attractive features. An efficient sampling scheme is developed to jointly generate the threshold values and the values of latent variable. With the simulated latent variable, the posterior distributions of other unknowns, for example, the number, locations, genetic effects, and genotypes of QTL, can be computed using existing algorithms for normally distributed traits. To this end, we provide a unified approach to mapping multiple QTL for continuous, binary, and ordinal traits. Utility and flexibility of the method are demonstrated using simulated data.

Algorithms↗

Skin segmentation using color pixel classification: analysis and comparison.

This paper presents a study of three important issues of the color pixel classification approach to skin segmentation: color representation, color quantization, and classification algorithm. Our analysis of several representative color spaces using the Bayesian classifier with the histogram technique shows that skin segmentation based on color pixel classification is largely unaffected by the choice of the color space. However, segmentation performance degrades when only chrominance channels are used in classification. Furthermore, we find that color quantization can be as low as 64 bins per channel, although higher histogram sizes give better segmentation performance. The Bayesian classifier with the histogram technique and the multilayer perceptron classifier are found to perform better compared to other tested classifiers, including three piecewise linear classifiers, three unimodal Gaussian classifiers, and a Gaussian mixture classifier.

Algorithms↗

Chromosome abnormalities in ovarian adenocarcinoma: III. Using breakpoint data to infer and test mathematical models for oncogenesis.

Cancer geneticists seek to identify genetic changes in tumor cells and to relate the genetic changes to tumor development. Because single changes can disrupt the cell cycle and promote other genetic changes, it is extremely hard to distinguish cause from effect. In this article we illustrate how 7 techniques from statistics, theoretical computer science, and phylogenetics can be used to infer and test possible models of tumor progression from single genome-wide descriptions of aberrations in a large sample of tumors. Specifically, we propose 4 tree models for tumor progression inferred from the large ovarian cancer data set described in the first 2 articles in this series. The models are derived from 2 different methods to select the non-random genetic aberrations and 2 different methods to infer the trees, given a set of events. Various aspects of the tree models are tested and extended by 5 methods: overall tests of independence, likelihood ratio tests, principal components analysis, directed acyclic graph modeling, and Bayesian survival analysis. All our methods lead to strikingly consistent conclusions about chromosomal breakpoints in ovarian adenocarcinoma, including (1) the non-random breakpoints in ovarian adenocarcinoma do not occur independently; (2) breakpoints in regions 1p3 and 11p1 are important early events and distinguish a class of tumors associated with poor prognosis; and (3) breakpoints in 1p1, 3p1, and 1q2 distinguish a class of ovarian tumors, and the breaks at 1p1 and 3p1 are associated with poor prognosis.

Adenocarcinoma↗

Data augmentation priors for Bayesian and semi-Bayes analyses of conditional-logistic and proportional-hazards regression.

Data augmentation priors have a long history in Bayesian data analysis. Formulae for such priors have been derived for generalized linear models, but their accuracy depends on two approximation steps. This note presents a method for using offsets as well as scaling factors to improve the accuracy of the approximations in logistic regression. This method produces an exceptionally simple form of data augmentation that allows it to be used with any standard package for conditional-logistic or proportional-hazards regression to perform Bayesian and semi-Bayes analyses of matched and survival data. The method is illustrated with an analysis of a matched case-control study of diet and breast cancer.

Algorithms↗

Chemical plume source localization.

This paper addresses the problem of estimating a likelihood map for the location of the source of a chemical plume using an autonomous vehicle as a sensor probe in a fluid flow. The fluid flow is assumed to have a high Reynolds number. Therefore, the dispersion of the chemical is dominated by turbulence, resulting in an intermittent chemical signal. The vehicle is capable of detecting above-threshold chemical concentration and sensing the fluid flow velocity at the vehicle location. This paper reviews instances of biological plume tracing and reviews previous strategies for a vehicle-based plume tracing. The main contribution is a new source-likelihood mapping approach based on Bayesian inference methods. Using this Bayesian methodology, the source-likelihood map is propagated through time and updated in response to both detection and nondetection events. Examples are included that use data from in-water testing to compare the mapping approach derived herein with the map derived using a previously existing technique.

Air Pollution↗

Bayesian estimation of cost-effectiveness ratios from clinical trials.

Estimation of the incremental cost-effectiveness ratio (ICER) is difficult for several reasons: treatments that decrease both cost and effectiveness and treatments that increase both cost and effectiveness can yield identical values of the ICER; the ICER is a discontinuous function of the mean difference in effectiveness; and the standard estimate of the ICER is a ratio. To address these difficulties, we have developed a Bayesian methodology that involves computing posterior probabilities for the four quadrants and separate interval estimates of ICER for the quadrants of interest. We compute these quantities by simulating draws from the posterior distribution of the cost and effectiveness parameters and tabulating the appropriate posterior probabilities and quantiles. We demonstrate the method by re-analysing three previously published clinical trials.

Bayes Theorem↗

The role of quantitative (18)F-FDG PET studies for the differentiation of malignant and benign bone lesions.

UNLABELLED: The role of quantitative (18)F-FDG PET studies for the differentiation of benign and malignant bone lesions is still an open question. METHODS: Our evaluation included 83 patients with 37 histologically proven malignancies and 46 benign lesions. Thirty-five of the 46 benign lesions were histologically confirmed. The (18)F-FDG studies were accomplished as a dynamic series for 60 min. Evaluation of the (18)F-FDG kinetics was performed using the following parameters: standardized uptake value (SUV), global influx (Ki), computation of the transport constants K1-k4 with consideration of the distribution volume (VB) according to a 2-tissue-compartment model, fractal dimension based on the box-counting procedure (parameter for the inhomogeneity of the tumors). RESULTS: The mean SUV, the vascular fraction VB, K1, and k3 were higher in malignant tumors compared with benign lesions (t test; P < 0.05). Although the (18)F-FDG SUV was helpful to differentiate benign and malignant tumors, there was some overlap, which limited the diagnostic accuracy. On the basis of the discriminant analysis, the SUV alone showed a sensitivity of only 54.05%, a specificity of 91.30%, and a diagnostic accuracy of 74.70%. The fractal dimension was superior and showed a sensitivity of 71.88%, a specificity of 81.58%, and an accuracy of 77.14%. The combination of SUV, fractal dimension, VB, K1-k4, and Ki revealed the best results with a sensitivity of 75.86%, a specificity of 97.22%, and an accuracy of 87.69%. Bayesian analysis showed true-positive results at the level of 0.8 for a low prevalence of disease (0.235) if the full kinetic data were used in the evaluation. CONCLUSION: (18)F-FDG PET has a high specificity for the exclusion of a malignant bone tumor. Evaluation of the full (18)F-FDG kinetics and the application of discriminant analysis are required and can be used prospectively to classify a bone lesion as malignant or benign.

Bayes Theorem↗

Bayesian model selection and averaging in additive and proportional hazards models.

Although Cox proportional hazards regression is the default analysis for time to event data, there is typically uncertainty about whether the effects of a predictor are more appropriately characterized by a multiplicative or additive model. To accommodate this uncertainty, we place a model selection prior on the coefficients in an additive-multiplicative hazards model. This prior assigns positive probability, not only to the model that has both additive and multiplicative effects for each predictor, but also to sub-models corresponding to no association, to only additive effects, and to only proportional effects. The additive component of the model is constrained to ensure non-negative hazards, a condition often violated by current methods. After augmenting the data with Poisson latent variables, the prior is conditionally conjugate, and posterior computation can proceed via an efficient Gibbs sampling algorithm. Simulation study results are presented, and the methodology is illustrated using data from the Framingham heart study.

Age of Onset↗

Cheminformatics analysis and learning in a data pipelining environment.

Workflow technology is being increasingly applied in discovery information to organize and analyze data. SciTegic's Pipeline Pilot is a chemically intelligent implementation of a workflow technology known as data pipelining. It allows scientists to construct and execute workflows using components that encapsulate many cheminformatics based algorithms. In this paper we review SciTegic's methodology for molecular fingerprints, molecular similarity, molecular clustering, maximal common subgraph search and Bayesian learning. Case studies are described showing the application of these methods to the analysis of discovery data such as chemical series and high throughput screening results. The paper demonstrates that the methods are well suited to a wide variety of tasks such as building and applying predictive models of screening data, identifying molecules for lead optimization and the organization of molecules into families with structural commonality.

Anti-Infective Agents↗