Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Quantifying uncertainty in geoacoustic inversion. I. A fast Gibbs sampler approach.

This paper develops a new approach to estimating seabed geoacoustic properties and their uncertainties based on a Bayesian formulation of matched-field inversion. In Bayesian inversion, the solution is characterized by its posterior probability density (PPD), which combines prior information about the model with information from an observed data set. To interpret the multi-dimensional PPD requires calculation of its moments, such as the mean, covariance, and marginal distributions, which provide parameter estimates and uncertainties. Computation of these moments involves estimating multi-dimensional integrals of the PPD, which is typically carried out using a sampling procedure. Important goals for an effective Bayesian algorithm are to obtain efficient, unbiased sampling of these moments, and to verify convergence of the sample. This is accomplished here using a Gibbs sampler (GS) approach based on the Metropolis algorithm, which also forms the basis for simulated annealing (SA). Although GS can be computationally slow in its basic form, just as modifications to SA have produced much faster optimization algorithms, the GS is modified here to produce an efficient algorithm referred to as the fast Gibbs sampler (FGS). An automated convergence criterion is employed based on monitoring the difference between two independent FGS samples collected in parallel. Comparison of FGS, GS, and Monte Carlo integration for noisy synthetic benchmark test cases indicates that FGS provides rigorous estimates of PPD moments while requiring orders of magnitude less computation time.

Acoustics↗

A Bayesian model of stereopsis depth and motion direction discrimination.

The extraction of stereoscopic depth from retinal disparity, and motion direction from two-frame kinematograms, requires the solution of a correspondence problem. In previous psychophysical work [Read and Eagle (2000) Vision Res 40: 3345-3358], we compared the performance of the human stereopsis and motion systems with correlated and anti-correlated stimuli. We found that, although the two systems performed similarly for narrow-band stimuli, broadband anti-correlated kinematograms produced a strong perception of reversed motion, whereas the stereograms appeared merely rivalrous. I now model these psychophysical data with a computational model of the correspondence problem based on the known properties of visual cortical cells. Noisy retinal images are filtered through a set of Fourier channels tuned to different spatial frequencies and orientations. Within each channel, a Bayesian analysis incorporating a prior preference for small disparities is used to assess the probability of each possible match. Finally, information from the different channels is combined to arrive at a judgement of stimulus disparity. Each model system--stereopsis and motion--has two free parameters: the amount of noise they are subject to, and the strength of their preference for small disparities. By adjusting these parameters independently for each system, qualitative matches are produced to psychophysical data, for both correlated and anti-correlated stimuli, across a range of spatial frequency and orientation bandwidths. The motion model is found to require much higher noise levels and a weaker preference for small disparities. This makes the motion model more tolerant of poor-quality reverse-direction false matches encountered with anti-correlated stimuli, matching the strong perception of reversed motion that humans experience with these stimuli. In contrast, the lower noise level and tighter prior preference used with the stereopsis model means that it performs close to chance with anti-correlated stimuli, in accordance with human psychophysics. Thus, the key features of the experimental data can be reproduced assuming that the motion system experiences more effective noise than the stereoscopy system and imposes a less stringent preference for small disparities.

Animals↗

Bayesian coestimation of phylogeny and sequence alignment.

BACKGROUND: Two central problems in computational biology are the determination of the alignment and phylogeny of a set of biological sequences. The traditional approach to this problem is to first build a multiple alignment of these sequences, followed by a phylogenetic reconstruction step based on this multiple alignment. However, alignment and phylogenetic inference are fundamentally interdependent, and ignoring this fact leads to biased and overconfident estimations. Whether the main interest be in sequence alignment or phylogeny, a major goal of computational biology is the co-estimation of both. RESULTS: We developed a fully Bayesian Markov chain Monte Carlo method for coestimating phylogeny and sequence alignment, under the Thorne-Kishino-Felsenstein model of substitution and single nucleotide insertion-deletion (indel) events. In our earlier work, we introduced a novel and efficient algorithm, termed the "indel peeling algorithm", which includes indels as phylogenetically informative evolutionary events, and resembles Felsenstein's peeling algorithm for substitutions on a phylogenetic tree. For a fixed alignment, our extension analytically integrates out both substitution and indel events within a proper statistical model, without the need for data augmentation at internal tree nodes, allowing for efficient sampling of tree topologies and edge lengths. To additionally sample multiple alignments, we here introduce an efficient partial Metropolized independence sampler for alignments, and combine these two algorithms into a fully Bayesian co-estimation procedure for the alignment and phylogeny problem. Our approach results in estimates for the posterior distribution of evolutionary rate parameters, for the maximum a-posteriori (MAP) phylogenetic tree, and for the posterior decoding alignment. Estimates for the evolutionary tree and multiple alignment are augmented with confidence estimates for each node height and alignment column. Our results indicate that the patterns in reliability broadly correspond to structural features of the proteins, and thus provides biologically meaningful information which is not existent in the usual point-estimate of the alignment. Our methods can handle input data of moderate size (10-20 protein sequences, each 100-200 bp), which we analyzed overnight on a standard 2 GHz personal computer. CONCLUSION: Joint analysis of multiple sequence alignment, evolutionary trees and additional evolutionary parameters can be now done within a single coherent statistical framework.

Algorithms↗

Hierarchical Bayesian modeling of spatially correlated health service outcome and utilization rates.

We present Bayesian hierarchical spatial models for spatially correlated small-area health service outcome and utilization rates, with a particular emphasis on the estimation of both measured and unmeasured or unknown covariate effects. This Bayesian hierarchical model framework enables simultaneous modeling of fixed covariate effects and random residual effects. The random effects are modeled via Bayesian prior specifications reflecting spatial heterogeneity globally and relative homogeneity among neighboring areas. The model inference is implemented using Markov chain Monte Carlo methods. Specifically, a hybrid Markov chain Monte Carlo algorithm (Neal, 1995, Bayesian Learning for Neural Networks; Gustafson, MacNab, and Wen, 2003, Statistics and Computing, to appear) is used for posterior sampling of the random effects. To illustrate relevant problems, methods, and techniques, we present an analysis of regional variation in intraventricular hemorrhage incidence rates among neonatal intensive care unit patients across Canada.

Bayes Theorem↗

Multiple association analysis via simulated annealing (MASSA).

SUMMARY: Genome-wide association studies are now technically feasible and likely to become a fundamental tool in unraveling the ultimate genetic basis of complex traits. However, new statistical and computational methods need to be developed to extract the maximum information in a realistic computing time. Here we propose a new method for multiple association analysis via simulated annealing that allows for epistasis and any number of markers. It consists of finding the model with lowest Bayesian information criterion using simulated annealing. The data are described by means of a mixed model and new alternative models are proposed using a set of rules, e.g. new sites can be added (or deleted), or new epistatic interactions can be included between existing genetic factors. The method is illustrated with simulated and real data. AVAILABILITY: An executable version of the program (MASSA) running under the Linux OS is freely available, together with documentation, at http://www.icrea.es/pag.asp?id=Miguel.Perez.

Algorithms↗

Computerized coding of injury narrative data from the National Health Interview Survey.

OBJECTIVE: To investigate the accuracy of a computerized method for classifying injury narratives into external-cause-of-injury and poisoning (E-code) categories. METHODS: This study used injury narratives and corresponding E-codes assigned by experts from the 1997 and 1998 US National Health Interview Survey (NHIS). A Fuzzy Bayesian model was used to assign injury descriptions to 13 E-code categories. Sensitivity, specificity and positive predictive value were measured by comparing the computer generated codes with E-code categories assigned by experts. RESULTS: The computer program correctly classified 4695 (82.7%) of the 5677 injury narratives when multiple words were included as keywords in the model. The use of multiple-word predictors compared with using single words alone improved both the sensitivity and specificity of the computer generated codes. The program is capable of identifying and filtering out cases that would benefit most from manual coding. For example, the program could be used to code the narrative if the maximum probability of a category given the keywords in the narrative was at least 0.9. If the maximum probability was lower than 0.9 (which will be the case for approximately 33% of the narratives) the case would be filtered out for manual review. CONCLUSIONS: A computer program based on Fuzzy Bayes logic is capable of accurately categorizing cause-of-injury codes from injury narratives. The capacity to filter out certain cases for manual coding improves the utility of this process.

Forms and Records Control↗

Incorporation of splice site probability models for non-canonical introns improves gene structure prediction in plants.

MOTIVATION: The vast majority of introns in protein-coding genes of higher eukaryotes have a GT dinucleotide at their 5'-terminus and an AG dinucleotide at their 3' end. About 1-2% of introns are non-canonical, with the most abundant subtype of non-canonical introns being characterized by GC and AG dinucleotides at their 5'- and 3'-termini, respectively. Most current gene prediction software, whether based on ab initio or spliced alignment approaches, does not include explicit models for non-canonical introns or may exclude their prediction altogether. With present amounts of genome and transcript data, it is now possible to apply statistical methodology to non-canonical splice site prediction. We pursued one such approach and describe the training and implementation of GC-donor splice site models for Arabidopsis and rice, with the goal of exploring whether specific modeling of non-canonical introns can enhance gene structure prediction accuracy. RESULTS: Our results indicate that the incorporation of non-canonical splice site models yields dramatic improvements in annotating genes containing GC-AG and AT-AC non-canonical introns. Comparison of models shows differences between monocot and dicot species, but also suggests GC intron-specific biases independent of taxonomic clade. We also present evidence that GC-AG introns occur preferentially in genes with atypically high exon counts. AVAILABILITY: Source code for the updated versions of GeneSeqer and SplicePredictor (distributed with the GeneSeqer code) isavailable at http://bioinformatics.iastate.edu/bioinformatics2go/gs/download.html. Web servers for Arabidopsis, rice and other plant species are accessible at http://www.plantgdb.org/PlantGDB-cgi/GeneSeqer/AtGDBgs.cgi, http://www.plantgdb.org/PlantGDB-cgi/GeneSeqer/OsGDBgs.cgi and http://www.plantgdb.org/PlantGDB-cgi/GeneSeqer/PlantGDBgs.cgi, respectively. A SplicePredictor web server is available at http://bioinformatics.iastate.edu/cgi-bin/sp.cgi. Software to generate training data and parameterizations for Bayesian splice site models is available at http://gremlin1.gdcb.iastate.edu/~volker/SB05B/BSSM4GSQ/

Base Composition↗

Bayesian pharmacokinetic estimation of vinorelbine in non-small-cell lung cancer patients.

OBJECTIVE: To develop a population pharmacokinetics of vinorelbine in a population of non-small-cell lung cancer (NSCLC) patients using a Bayesian estimation in order to calculate for any further patient, individual pharmacokinetic parameters from few blood samples. METHODS: Vinorelbine was given by a 15-min infusion (30 mg x m(-2)) to eight patients with NSCLC. Its serum concentration was determined by HPLC and its pharmacokinetics was described by a three-compartment open model with elimination from the central compartment. Volume of the central compartment (V1) and rate constants (k10, k12, k21, k13, k31) were selected as population pharmacokinetic parameters and computed by non-linear regression (two-step approach) from 14 to 18 concentration measurements per course. Subsequently, these parameters were used by the Bayesian estimator to calculate individual pharmacokinetics from only 2 or 3 measured concentrations. RESULTS: The population mean values (CV%) of V1, k10, k12, k21, k13, k31, CL, t1/2gamma were respectively 21 l (55%), 3.2 h(-1) (29%), 7.7 h(-1) (74%), 1.3 h(-1) (67%), 4.7 h(-1) (53%), 0.04 h(-1) (20%), 57 l x h(-1) (31%) and 43 h (36%). The comparison of results obtained from the Bayesian estimator and from the three-compartment model showed that CL and t1/2gamma were well predicted (relative deviation: +/- 12 to 22%) by the Bayesian method using only two blood samples. CONCLUSION: We demonstrated that Bayesian estimation allows, at minimal cost and minimal disturbance for the patient, the determination of several vinorelbine pharmacokinetic parameters and therefore dose adaptation from as few as two drug concentrations, measured at 6 h and 24 h after infusion.

Aged↗

Open-loop-feedback control of serum drug concentrations: pharmacokinetic approaches to drug therapy.

Recent developments to optimize open-loop-feedback control of drug dosage regimens, generally applicable to pharmacokinetically oriented therapy with many drugs, involve computation of patient-individualized strategies for obtaining desired serum drug concentrations. Analyses of past therapy are performed by least squares, extended least squares, and maximum a posteriori probability Bayesian methods of fitting pharmacokinetic models to serum level data. Future possibilities for truly optimal open-loop-feedback therapy with full Bayesian methods, and conceivably for optimal closed-loop therapy in such data-poor clinical situations, are also discussed. Implementation of these various therapeutic strategies, using automated, locally controlled infusion devices, has also been achieved in prototype form.

Aged↗

Dynamics of the evolution of learning algorithms by selection.

We study the evolution of artificial learning systems by means of selection. Genetic programming is used to generate populations of programs that implement algorithms used by neural network classifiers to learn a rule in a supervised learning scenario. In contrast to concentrating on final results, which would be the natural aim while designing good learning algorithms, we study the evolution process. Phenotypic and genotypic entropies, which describe the distribution of fitness and of symbols, respectively, are used to monitor the dynamics. We identify significant functional structures responsible for the improvements in the learning process. In particular, some combinations of variables and operators are useful in assessing performance in rule extraction and can thus implement annealing of the learning schedule. We also find combinations that can signal surprise, measured on a single example, by the difference between predicted and correct classification. When such favorable structures appear, they are disseminated on very short time scales throughout the population. Due to such abruptness they can be thought of as dynamical transitions. But foremost, we find a strict temporal order of such discoveries. Structures that measure performance are never useful before those for measuring surprise. Invasions of the population by such structures in the reverse order were never observed. Asymptotically, the generalization ability approaches Bayesian results.

Algorithms↗

Bayesian inference for a generalized population attributable fraction: the impact of early vitamin A levels on chronic lung disease in very low birthweight infants.

In this paper, the population attributable fraction is studied using the potential responses framework of Rubin's causal model. This framework facilitates definition of a general measure of population attributable effect which can accommodate many-valued and multivariate exposures as well as many-valued responses. Inferential issues are considered from the Bayesian perspective. Finite population inference is emphasized with inference in the case of a fully observed population given particular attention. The key inferential issue concerns computation of the posterior distribution of unobserved potential responses, given observed responses, exposures and covariates. A dependency on model parameters about which observed data are uninformative is highlighted and this reflects the unobservable nature of causal effects. In an application to a small cohort study of respiratory problems in very low birthweight infants, posterior inferences were found to be insensitive to assumptions concerning the joint distribution of potential response variables but sensitive to the assumption of weak ignorability, a weaker form of the more familiar assumption of no confounding by omitted covariates. In a model-based set-up, the weak ignorability assumption is identified with setting a model parameter to zero, and consequently uncertainty concerning this assumption can, in principle, be handled via the prior distribution for the model parameters.

Bayes Theorem↗

A rapid computational filter for cytochrome P450 1A2 inhibition potential of compound libraries.

QSAR models for a diverse set of compounds for cytochrome P450 1A2 inhibition have been produced using 4 statistical approaches; partial least squares (PLS), multiple linear regression (MLR), classification and regression trees (CART), and bayesian neural networks (BNN). The models complement one another and have identified the following descriptors as important features for CYP1A2 inhibition; lipophilicity, aromaticity, charge, and the HOMO/LUMO energies. Furthermore all models are global and have been used to predict a diverse independent set of compounds. For the first time in the field of QSAR, the kappa index of agreement has comprehensively been used to assess the overall accuracy of the model's predictive power. The models are statistically significant and can be used as a rapid computational filter for cytochrome P450 1A2 inhibition potential of compound libraries.

Bayes Theorem↗

Triple-goal estimates for disease mapping.

Maps of regional morbidity and mortality rates play an important role in assessing environmental equity. They provide effective tools for identifying areas with potentially elevated risk, determining spatial trend, and formulating and validating aetiological hypotheses about disease. Bayes and empirical Bayes methods produce stable small-area rate estimates that retain geographic and demographic resolution. The beauty of the Bayesian approach lies in its ability to structure complicated models, inferential goals and analyses. Three inferential goals are relevant to disease mapping and risk assessment: (i) computing accurate estimates of disease rates in small geographic areas; (ii) estimating the distribution of disease rates over the region; (iii) ranking the disease rates so that environmental investigation can be prioritized. No single set of estimates can simultaneously optimize these three goals, and Shen and Louis propose a set of estimates that perform well on all three goals. These are optimal for estimating the distribution of rates and for ranking, and maintain a high accuracy in estimating area-specific rates. However, the Shen/Louis method is sensitive to choice of priors. To address this issue we introduce a robustified version of the method based on a smoothed non-parametric estimate of the prior. We evaluate the performance of this method through a simulation study, and illustrate it using a data set of county-specific lung cancer rates in Ohio.

Algorithms↗

Population toxicokinetics of tetrachloroethylene.

In assessing the distribution and metabolism of toxic compounds in the body, measurements are not always feasible for ethical or technical reasons. Computer modeling offers a reasonable alternative, but the variability and complexity of biological systems pose unique challenges in model building and adjustment. Recent tools from population pharmacokinetics, Bayesian statistical inference, and physiological modeling can be brought together to solve these problems. As an example, we modeled the distribution and metabolism of tetrachloroethylene (PERC) in humans. We derive statistical distributions for the parameters of a physiological model of PERC, on the basis of data from Monster et al. (1979). The model adequately fits both prior physiological information and experimental data. An estimate of the relationship between PERC exposure and fraction metabolized is obtained. Our median population estimate for the fraction of inhaled tetrachloroethylene that is metabolized, at exposure levels exceeding current occupational standards, is 1.5% [95% confidence interval (0.52%, 4.1%)]. At levels approaching ambient inhalation exposure (0.001 ppm), the median estimate of the fraction metabolized is much higher, at 36% [95% confidence interval (15%, 58%)]. This disproportionality should be taken into account when deriving safe exposure limits for tetrachloroethylene and deserves to be verified by further experiments.

Administration, Inhalation↗

The pulmonologist's perspective regarding the solitary pulmonary nodule.

The pulmonologist's goal in managing a patient with a solitary pulmonary nodule is to distinguish the benign from malignant nodule and, where malignancy is either confirmed or strongly suspected, to expedite resection. By using established clinical features (eg, age, smoking status) and radiographic findings (eg, calcification, growth rate, size), a probability of malignancy can be determined. If necessary, noninvasive or adjuvant invasive testing is used to alter the probability to one that permits observation or demands resection. The proper use of these tests mandates knowledge about their performance characteristics. Decision-analytic approaches, using Bayesian analysis, may assist with the calculation of probability. These models have not consistently outperformed the clinician or adjuvant testing. The use of low-dose computed tomography (CT) scanning as a screening tool has led to the discovery of many small, indeterminate nodules. Management decisions for these nodules are influenced by their low prevalence of malignancy and small size. Future advances will add to our ability to effectively meet our stated goal.

Decision Making↗

A Bayesian approach to DNA sequence segmentation.

Many deoxyribonucleic acid (DNA) sequences display compositional heterogeneity in the form of segments of similar structure. This article describes a Bayesian method that identifies such segments by using a Markov chain governed by a hidden Markov model. Markov chain Monte Carlo (MCMC) techniques are employed to compute all posterior quantities of interest and, in particular, allow inferences to be made regarding the number of segment types and the order of Markov dependence in the DNA sequence. The method is applied to the segmentation of the bacteriophage lambda genome, a common benchmark sequence used for the comparison of statistical segmentation algorithms.

Algorithms↗

Causal protein-signaling networks derived from multiparameter single-cell data.

Machine learning was applied for the automated derivation of causal influences in cellular signaling networks. This derivation relied on the simultaneous measurement of multiple phosphorylated protein and phospholipid components in thousands of individual primary human immune system cells. Perturbing these cells with molecular interventions drove the ordering of connections between pathway components, wherein Bayesian network computational methods automatically elucidated most of the traditionally reported signaling relationships and predicted novel interpathway network causalities, which we verified experimentally. Reconstruction of network models from physiologically relevant primary single cells might be applied to understanding native-state tissue signaling biology, complex drug actions, and dysfunctional signaling in diseased cells.

Algorithms↗

Performance comparison of two-point linkage methods using microsatellite markers flanking known disease locations.

The Genetic Analysis Workshop 14 simulated data presents an interesting, challenging, and plausible example of a complex disease interaction in a dataset. This paper summarizes the ease of detection for each of the simulated Kofendrerd Personality Disorder (KPD) genes across all of the replicates for five standard linkage statistics. Using the KPD affection status, we have analyzed the microsatellite markers flanking each of the disease genes, plus an additional 2 markers that were not linked to any of the disease loci. All markers were analyzed using the following two-point linkage methods: 1) a MMLS, which is a standard admixture LOD score maximized over theta, alpha, and mode of inheritance, 2) a MLS calculated by GENEHUNTER, 3) the Kong and Cox LOD score as computed by MERLIN, 4) a MOD score (standard heterogeneity LOD maximized over theta, alpha, and a grid of genetic model parameters), and 5) the PPL, a Bayesian statistic that directly measures the strength of evidence for linkage to a marker. All of the major loci (D1-D4) were detectable with varying probabilities in the different populations. However, the modifier genes (D5 and D6) were difficult to detect, with similar distributions under the null and alternative across populations and statistics. The pooling of the four datasets in each replicate (n = 350 pedigrees) greatly improved the chance of detecting the major genes using all five methods, but failed to increase the chance to detect D5 and D6.

Chromosome Mapping↗