Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Using Bayesian networks to analyze expression data.

DNA hybridization arrays simultaneously measure the expression level for thousands of genes. These measurements provide a "snapshot" of transcription levels within the cell. A major challenge in computational biology is to uncover, from such measurements, gene/protein interactions and key biological features of cellular systems. In this paper, we propose a new framework for discovering interactions between genes based on multiple expression measurements. This framework builds on the use of Bayesian networks for representing statistical dependencies. A Bayesian network is a graph-based model of joint multivariate probability distributions that captures properties of conditional independence between variables. Such models are attractive for their ability to describe complex stochastic processes and because they provide a clear methodology for learning from (noisy) observations. We start by showing how Bayesian networks can describe interactions between genes. We then describe a method for recovering gene interactions from microarray data using tools for learning Bayesian networks. Finally, we demonstrate this method on the S. cerevisiae cell-cycle measurements of Spellman et al. (1998).

Algorithms↗

Prediction of caspase cleavage sites using Bayesian bio-basis function neural networks.

MOTIVATION: Apoptosis has drawn the attention of researchers because of its importance in treating some diseases through finding a proper way to block or slow down the apoptosis process. Having understood that caspase cleavage is the key to apoptosis, we find novel methods or algorithms are essential for studying the specificity of caspase cleavage activity and this helps the effective drug design. As bio-basis function neural networks have proven to outperform some conventional neural learning algorithms, there is a motivation, in this study, to investigate the application of bio-basis function neural networks for the prediction of caspase cleavage sites. RESULTS: Thirteen protein sequences with experimentally determined caspase cleavage sites were downloaded from NCBI. Bayesian bio-basis function neural networks are investigated and the comparisons with single-layer perceptrons, multilayer perceptrons, the original bio-basis function neural networks and support vector machines are given. The impact of the sliding window size used to generate sub-sequences for modelling on prediction accuracy is studied. The results show that the Bayesian bio-basis function neural network with two Gaussian distributions for model parameters (weights) performed the best and the highest prediction accuracy is 97.15 +/- 1.13%. AVAILABILITY: The package of Bayesian bio-basis function neural network can be obtained by request to the author.

Algorithms↗

Inferring the location and effect of tumor suppressor genes by instability-selection modeling of allelic-loss data.

Cancerous tumor growth creates cells with abnormal DNA. Allelic-loss experiments identify genomic deletions in cancer cells, but sources of variation and intrinsic dependencies complicate inference about the location and effect of suppressor genes; such genes are the target of these experiments and are thought to be involved in tumor development. We investigate properties of an instability-selection model of allelic-loss data, including likelihood-based parameter estimation and hypothesis testing. By considering a special complete-data case, we derive an approximate calibration method for hypothesis tests of sporadic deletion. Parametric bootstrap and Bayesian computations are also developed. Data from three allelic-loss studies are reanalyzed to illustrate the methods.

Adenocarcinoma↗

Bayesian model of Snellen visual acuity.

A Bayesian model of Snellen visual acuity (VA) has been developed that, as far as we know, is the first one that includes the three main stages of VA: (1) optical degradations, (2) neural image representation and contrast thresholding, and (3) character recognition. The retinal image of a Snellen test chart is obtained from experimental wave-aberration data. Then a subband image decomposition with a set of visual channels tuned to different spatial frequencies and orientations is applied to the retinal image, as in standard computational models of early cortical image representation. A neural threshold is applied to the contrast responses to include the effect of the neural contrast sensitivity. The resulting image representation is the base of a Bayesian pattern-recognition method robust to the presence of optical aberrations. The model is applied to images containing sets of letter optotypes at different scales, and the number of correct answers is obtained at each scale; the final output is the decimal Snellen VA. The model has no free parameters to adjust. The main input data are the eye's optical aberrations, and standard values are used for all other parameters, including the Stiles-Crawford effect, visual channels, and neural contrast threshold, when no subject specific values are available. When aberrations are large, Snellen VA involving pattern recognition differs from grating acuity, which is based on a simpler detection (or orientation-discrimination) task and hence is basically unaffected by phase distortions introduced by the optical transfer function. A preliminary test of the model in one subject produced close agreement between actual measurements and predicted VA values. Two examples are also included: (1) application of the method to the prediction of the VAin refractive-surgery patients and (2) simulation of the VA attainable by correcting ocular aberrations.

Bayes Theorem↗

Quantifying uncertainty in geoacoustic inversion. I. A fast Gibbs sampler approach.

This paper develops a new approach to estimating seabed geoacoustic properties and their uncertainties based on a Bayesian formulation of matched-field inversion. In Bayesian inversion, the solution is characterized by its posterior probability density (PPD), which combines prior information about the model with information from an observed data set. To interpret the multi-dimensional PPD requires calculation of its moments, such as the mean, covariance, and marginal distributions, which provide parameter estimates and uncertainties. Computation of these moments involves estimating multi-dimensional integrals of the PPD, which is typically carried out using a sampling procedure. Important goals for an effective Bayesian algorithm are to obtain efficient, unbiased sampling of these moments, and to verify convergence of the sample. This is accomplished here using a Gibbs sampler (GS) approach based on the Metropolis algorithm, which also forms the basis for simulated annealing (SA). Although GS can be computationally slow in its basic form, just as modifications to SA have produced much faster optimization algorithms, the GS is modified here to produce an efficient algorithm referred to as the fast Gibbs sampler (FGS). An automated convergence criterion is employed based on monitoring the difference between two independent FGS samples collected in parallel. Comparison of FGS, GS, and Monte Carlo integration for noisy synthetic benchmark test cases indicates that FGS provides rigorous estimates of PPD moments while requiring orders of magnitude less computation time.

Acoustics↗

A Bayesian model of stereopsis depth and motion direction discrimination.

The extraction of stereoscopic depth from retinal disparity, and motion direction from two-frame kinematograms, requires the solution of a correspondence problem. In previous psychophysical work [Read and Eagle (2000) Vision Res 40: 3345-3358], we compared the performance of the human stereopsis and motion systems with correlated and anti-correlated stimuli. We found that, although the two systems performed similarly for narrow-band stimuli, broadband anti-correlated kinematograms produced a strong perception of reversed motion, whereas the stereograms appeared merely rivalrous. I now model these psychophysical data with a computational model of the correspondence problem based on the known properties of visual cortical cells. Noisy retinal images are filtered through a set of Fourier channels tuned to different spatial frequencies and orientations. Within each channel, a Bayesian analysis incorporating a prior preference for small disparities is used to assess the probability of each possible match. Finally, information from the different channels is combined to arrive at a judgement of stimulus disparity. Each model system--stereopsis and motion--has two free parameters: the amount of noise they are subject to, and the strength of their preference for small disparities. By adjusting these parameters independently for each system, qualitative matches are produced to psychophysical data, for both correlated and anti-correlated stimuli, across a range of spatial frequency and orientation bandwidths. The motion model is found to require much higher noise levels and a weaker preference for small disparities. This makes the motion model more tolerant of poor-quality reverse-direction false matches encountered with anti-correlated stimuli, matching the strong perception of reversed motion that humans experience with these stimuli. In contrast, the lower noise level and tighter prior preference used with the stereopsis model means that it performs close to chance with anti-correlated stimuli, in accordance with human psychophysics. Thus, the key features of the experimental data can be reproduced assuming that the motion system experiences more effective noise than the stereoscopy system and imposes a less stringent preference for small disparities.

Animals↗

Bayesian coestimation of phylogeny and sequence alignment.

BACKGROUND: Two central problems in computational biology are the determination of the alignment and phylogeny of a set of biological sequences. The traditional approach to this problem is to first build a multiple alignment of these sequences, followed by a phylogenetic reconstruction step based on this multiple alignment. However, alignment and phylogenetic inference are fundamentally interdependent, and ignoring this fact leads to biased and overconfident estimations. Whether the main interest be in sequence alignment or phylogeny, a major goal of computational biology is the co-estimation of both. RESULTS: We developed a fully Bayesian Markov chain Monte Carlo method for coestimating phylogeny and sequence alignment, under the Thorne-Kishino-Felsenstein model of substitution and single nucleotide insertion-deletion (indel) events. In our earlier work, we introduced a novel and efficient algorithm, termed the "indel peeling algorithm", which includes indels as phylogenetically informative evolutionary events, and resembles Felsenstein's peeling algorithm for substitutions on a phylogenetic tree. For a fixed alignment, our extension analytically integrates out both substitution and indel events within a proper statistical model, without the need for data augmentation at internal tree nodes, allowing for efficient sampling of tree topologies and edge lengths. To additionally sample multiple alignments, we here introduce an efficient partial Metropolized independence sampler for alignments, and combine these two algorithms into a fully Bayesian co-estimation procedure for the alignment and phylogeny problem. Our approach results in estimates for the posterior distribution of evolutionary rate parameters, for the maximum a-posteriori (MAP) phylogenetic tree, and for the posterior decoding alignment. Estimates for the evolutionary tree and multiple alignment are augmented with confidence estimates for each node height and alignment column. Our results indicate that the patterns in reliability broadly correspond to structural features of the proteins, and thus provides biologically meaningful information which is not existent in the usual point-estimate of the alignment. Our methods can handle input data of moderate size (10-20 protein sequences, each 100-200 bp), which we analyzed overnight on a standard 2 GHz personal computer. CONCLUSION: Joint analysis of multiple sequence alignment, evolutionary trees and additional evolutionary parameters can be now done within a single coherent statistical framework.

Algorithms↗

Hierarchical Bayesian modeling of spatially correlated health service outcome and utilization rates.

We present Bayesian hierarchical spatial models for spatially correlated small-area health service outcome and utilization rates, with a particular emphasis on the estimation of both measured and unmeasured or unknown covariate effects. This Bayesian hierarchical model framework enables simultaneous modeling of fixed covariate effects and random residual effects. The random effects are modeled via Bayesian prior specifications reflecting spatial heterogeneity globally and relative homogeneity among neighboring areas. The model inference is implemented using Markov chain Monte Carlo methods. Specifically, a hybrid Markov chain Monte Carlo algorithm (Neal, 1995, Bayesian Learning for Neural Networks; Gustafson, MacNab, and Wen, 2003, Statistics and Computing, to appear) is used for posterior sampling of the random effects. To illustrate relevant problems, methods, and techniques, we present an analysis of regional variation in intraventricular hemorrhage incidence rates among neonatal intensive care unit patients across Canada.

Bayes Theorem↗

Multiple association analysis via simulated annealing (MASSA).

SUMMARY: Genome-wide association studies are now technically feasible and likely to become a fundamental tool in unraveling the ultimate genetic basis of complex traits. However, new statistical and computational methods need to be developed to extract the maximum information in a realistic computing time. Here we propose a new method for multiple association analysis via simulated annealing that allows for epistasis and any number of markers. It consists of finding the model with lowest Bayesian information criterion using simulated annealing. The data are described by means of a mixed model and new alternative models are proposed using a set of rules, e.g. new sites can be added (or deleted), or new epistatic interactions can be included between existing genetic factors. The method is illustrated with simulated and real data. AVAILABILITY: An executable version of the program (MASSA) running under the Linux OS is freely available, together with documentation, at http://www.icrea.es/pag.asp?id=Miguel.Perez.

Algorithms↗

Computerized coding of injury narrative data from the National Health Interview Survey.

OBJECTIVE: To investigate the accuracy of a computerized method for classifying injury narratives into external-cause-of-injury and poisoning (E-code) categories. METHODS: This study used injury narratives and corresponding E-codes assigned by experts from the 1997 and 1998 US National Health Interview Survey (NHIS). A Fuzzy Bayesian model was used to assign injury descriptions to 13 E-code categories. Sensitivity, specificity and positive predictive value were measured by comparing the computer generated codes with E-code categories assigned by experts. RESULTS: The computer program correctly classified 4695 (82.7%) of the 5677 injury narratives when multiple words were included as keywords in the model. The use of multiple-word predictors compared with using single words alone improved both the sensitivity and specificity of the computer generated codes. The program is capable of identifying and filtering out cases that would benefit most from manual coding. For example, the program could be used to code the narrative if the maximum probability of a category given the keywords in the narrative was at least 0.9. If the maximum probability was lower than 0.9 (which will be the case for approximately 33% of the narratives) the case would be filtered out for manual review. CONCLUSIONS: A computer program based on Fuzzy Bayes logic is capable of accurately categorizing cause-of-injury codes from injury narratives. The capacity to filter out certain cases for manual coding improves the utility of this process.

Forms and Records Control↗

Incorporation of splice site probability models for non-canonical introns improves gene structure prediction in plants.

MOTIVATION: The vast majority of introns in protein-coding genes of higher eukaryotes have a GT dinucleotide at their 5'-terminus and an AG dinucleotide at their 3' end. About 1-2% of introns are non-canonical, with the most abundant subtype of non-canonical introns being characterized by GC and AG dinucleotides at their 5'- and 3'-termini, respectively. Most current gene prediction software, whether based on ab initio or spliced alignment approaches, does not include explicit models for non-canonical introns or may exclude their prediction altogether. With present amounts of genome and transcript data, it is now possible to apply statistical methodology to non-canonical splice site prediction. We pursued one such approach and describe the training and implementation of GC-donor splice site models for Arabidopsis and rice, with the goal of exploring whether specific modeling of non-canonical introns can enhance gene structure prediction accuracy. RESULTS: Our results indicate that the incorporation of non-canonical splice site models yields dramatic improvements in annotating genes containing GC-AG and AT-AC non-canonical introns. Comparison of models shows differences between monocot and dicot species, but also suggests GC intron-specific biases independent of taxonomic clade. We also present evidence that GC-AG introns occur preferentially in genes with atypically high exon counts. AVAILABILITY: Source code for the updated versions of GeneSeqer and SplicePredictor (distributed with the GeneSeqer code) isavailable at http://bioinformatics.iastate.edu/bioinformatics2go/gs/download.html. Web servers for Arabidopsis, rice and other plant species are accessible at http://www.plantgdb.org/PlantGDB-cgi/GeneSeqer/AtGDBgs.cgi, http://www.plantgdb.org/PlantGDB-cgi/GeneSeqer/OsGDBgs.cgi and http://www.plantgdb.org/PlantGDB-cgi/GeneSeqer/PlantGDBgs.cgi, respectively. A SplicePredictor web server is available at http://bioinformatics.iastate.edu/cgi-bin/sp.cgi. Software to generate training data and parameterizations for Bayesian splice site models is available at http://gremlin1.gdcb.iastate.edu/~volker/SB05B/BSSM4GSQ/

Base Composition↗

A hierarchical Naïve Bayes Model for handling sample heterogeneity in classification problems: an application to tissue microarrays.

BACKGROUND: Uncertainty often affects molecular biology experiments and data for different reasons. Heterogeneity of gene or protein expression within the same tumor tissue is an example of biological uncertainty which should be taken into account when molecular markers are used in decision making. Tissue Microarray (TMA) experiments allow for large scale profiling of tissue biopsies, investigating protein patterns characterizing specific disease states. TMA studies deal with multiple sampling of the same patient, and therefore with multiple measurements of same protein target, to account for possible biological heterogeneity. The aim of this paper is to provide and validate a classification model taking into consideration the uncertainty associated with measuring replicate samples. RESULTS: We propose an extension of the well-known Naïve Bayes classifier, which accounts for biological heterogeneity in a probabilistic framework, relying on Bayesian hierarchical models. The model, which can be efficiently learned from the training dataset, exploits a closed-form of classification equation, thus providing no additional computational cost with respect to the standard Naïve Bayes classifier. We validated the approach on several simulated datasets comparing its performances with the Naïve Bayes classifier. Moreover, we demonstrated that explicitly dealing with heterogeneity can improve classification accuracy on a TMA prostate cancer dataset. CONCLUSION: The proposed Hierarchical Naïve Bayes classifier can be conveniently applied in problems where within sample heterogeneity must be taken into account, such as TMA experiments and biological contexts where several measurements (replicates) are available for the same biological sample. The performance of the new approach is better than the standard Naïve Bayes model, in particular when the within sample heterogeneity is different in the different classes.

Algorithms↗

Bayesian pharmacokinetic estimation of vinorelbine in non-small-cell lung cancer patients.

OBJECTIVE: To develop a population pharmacokinetics of vinorelbine in a population of non-small-cell lung cancer (NSCLC) patients using a Bayesian estimation in order to calculate for any further patient, individual pharmacokinetic parameters from few blood samples. METHODS: Vinorelbine was given by a 15-min infusion (30 mg x m(-2)) to eight patients with NSCLC. Its serum concentration was determined by HPLC and its pharmacokinetics was described by a three-compartment open model with elimination from the central compartment. Volume of the central compartment (V1) and rate constants (k10, k12, k21, k13, k31) were selected as population pharmacokinetic parameters and computed by non-linear regression (two-step approach) from 14 to 18 concentration measurements per course. Subsequently, these parameters were used by the Bayesian estimator to calculate individual pharmacokinetics from only 2 or 3 measured concentrations. RESULTS: The population mean values (CV%) of V1, k10, k12, k21, k13, k31, CL, t1/2gamma were respectively 21 l (55%), 3.2 h(-1) (29%), 7.7 h(-1) (74%), 1.3 h(-1) (67%), 4.7 h(-1) (53%), 0.04 h(-1) (20%), 57 l x h(-1) (31%) and 43 h (36%). The comparison of results obtained from the Bayesian estimator and from the three-compartment model showed that CL and t1/2gamma were well predicted (relative deviation: +/- 12 to 22%) by the Bayesian method using only two blood samples. CONCLUSION: We demonstrated that Bayesian estimation allows, at minimal cost and minimal disturbance for the patient, the determination of several vinorelbine pharmacokinetic parameters and therefore dose adaptation from as few as two drug concentrations, measured at 6 h and 24 h after infusion.

Aged↗

Open-loop-feedback control of serum drug concentrations: pharmacokinetic approaches to drug therapy.

Recent developments to optimize open-loop-feedback control of drug dosage regimens, generally applicable to pharmacokinetically oriented therapy with many drugs, involve computation of patient-individualized strategies for obtaining desired serum drug concentrations. Analyses of past therapy are performed by least squares, extended least squares, and maximum a posteriori probability Bayesian methods of fitting pharmacokinetic models to serum level data. Future possibilities for truly optimal open-loop-feedback therapy with full Bayesian methods, and conceivably for optimal closed-loop therapy in such data-poor clinical situations, are also discussed. Implementation of these various therapeutic strategies, using automated, locally controlled infusion devices, has also been achieved in prototype form.

Aged↗

Dynamics of the evolution of learning algorithms by selection.

We study the evolution of artificial learning systems by means of selection. Genetic programming is used to generate populations of programs that implement algorithms used by neural network classifiers to learn a rule in a supervised learning scenario. In contrast to concentrating on final results, which would be the natural aim while designing good learning algorithms, we study the evolution process. Phenotypic and genotypic entropies, which describe the distribution of fitness and of symbols, respectively, are used to monitor the dynamics. We identify significant functional structures responsible for the improvements in the learning process. In particular, some combinations of variables and operators are useful in assessing performance in rule extraction and can thus implement annealing of the learning schedule. We also find combinations that can signal surprise, measured on a single example, by the difference between predicted and correct classification. When such favorable structures appear, they are disseminated on very short time scales throughout the population. Due to such abruptness they can be thought of as dynamical transitions. But foremost, we find a strict temporal order of such discoveries. Structures that measure performance are never useful before those for measuring surprise. Invasions of the population by such structures in the reverse order were never observed. Asymptotically, the generalization ability approaches Bayesian results.

Algorithms↗

Bayesian inference for a generalized population attributable fraction: the impact of early vitamin A levels on chronic lung disease in very low birthweight infants.

In this paper, the population attributable fraction is studied using the potential responses framework of Rubin's causal model. This framework facilitates definition of a general measure of population attributable effect which can accommodate many-valued and multivariate exposures as well as many-valued responses. Inferential issues are considered from the Bayesian perspective. Finite population inference is emphasized with inference in the case of a fully observed population given particular attention. The key inferential issue concerns computation of the posterior distribution of unobserved potential responses, given observed responses, exposures and covariates. A dependency on model parameters about which observed data are uninformative is highlighted and this reflects the unobservable nature of causal effects. In an application to a small cohort study of respiratory problems in very low birthweight infants, posterior inferences were found to be insensitive to assumptions concerning the joint distribution of potential response variables but sensitive to the assumption of weak ignorability, a weaker form of the more familiar assumption of no confounding by omitted covariates. In a model-based set-up, the weak ignorability assumption is identified with setting a model parameter to zero, and consequently uncertainty concerning this assumption can, in principle, be handled via the prior distribution for the model parameters.

Bayes Theorem↗

A rapid computational filter for cytochrome P450 1A2 inhibition potential of compound libraries.

QSAR models for a diverse set of compounds for cytochrome P450 1A2 inhibition have been produced using 4 statistical approaches; partial least squares (PLS), multiple linear regression (MLR), classification and regression trees (CART), and bayesian neural networks (BNN). The models complement one another and have identified the following descriptors as important features for CYP1A2 inhibition; lipophilicity, aromaticity, charge, and the HOMO/LUMO energies. Furthermore all models are global and have been used to predict a diverse independent set of compounds. For the first time in the field of QSAR, the kappa index of agreement has comprehensively been used to assess the overall accuracy of the model's predictive power. The models are statistically significant and can be used as a rapid computational filter for cytochrome P450 1A2 inhibition potential of compound libraries.

Bayes Theorem↗

Triple-goal estimates for disease mapping.

Maps of regional morbidity and mortality rates play an important role in assessing environmental equity. They provide effective tools for identifying areas with potentially elevated risk, determining spatial trend, and formulating and validating aetiological hypotheses about disease. Bayes and empirical Bayes methods produce stable small-area rate estimates that retain geographic and demographic resolution. The beauty of the Bayesian approach lies in its ability to structure complicated models, inferential goals and analyses. Three inferential goals are relevant to disease mapping and risk assessment: (i) computing accurate estimates of disease rates in small geographic areas; (ii) estimating the distribution of disease rates over the region; (iii) ranking the disease rates so that environmental investigation can be prioritized. No single set of estimates can simultaneously optimize these three goals, and Shen and Louis propose a set of estimates that perform well on all three goals. These are optimal for estimating the distribution of rates and for ranking, and maintain a high accuracy in estimating area-specific rates. However, the Shen/Louis method is sensitive to choice of priors. To address this issue we introduce a robustified version of the method based on a smoothed non-parametric estimate of the prior. We evaluate the performance of this method through a simulation study, and illustrate it using a data set of county-specific lung cancer rates in Ohio.

Algorithms↗