Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Subcellular localization of the yeast proteome.

Protein localization data are a valuable information resource helpful in elucidating eukaryotic protein function. Here, we report the first proteome-scale analysis of protein localization within any eukaryote. Using directed topoisomerase I-mediated cloning strategies and genome-wide transposon mutagenesis, we have epitope-tagged 60% of the Saccharomyces cerevisiae proteome. By high-throughput immunolocalization of tagged gene products, we have determined the subcellular localization of 2744 yeast proteins. Extrapolating these data through a computational algorithm employing Bayesian formalism, we define the yeast localizome (the subcellular distribution of all 6100 yeast proteins). We estimate the yeast proteome to encompass approximately 5100 soluble proteins and >1000 transmembrane proteins. Our results indicate that 47% of yeast proteins are cytoplasmic, 13% mitochondrial, 13% exocytic (including proteins of the endoplasmic reticulum and secretory vesicles), and 27% nuclear/nucleolar. A subset of nuclear proteins was further analyzed by immunolocalization using surface-spread preparations of meiotic chromosomes. Of these proteins, 38% were found associated with chromosomal DNA. As determined from phenotypic analyses of nuclear proteins, 34% are essential for spore viability--a percentage nearly twice as great as that observed for the proteome as a whole. In total, this study presents experimentally derived localization data for 955 proteins of previously unknown function: nearly half of all functionally uncharacterized proteins in yeast. To facilitate access to these data, we provide a searchable database featuring 2900 fluorescent micrographs at http://ygac.med.yale.edu.

Algorithms↗

Bayesian inference of recent migration rates using multilocus genotypes.

A new Bayesian method that uses individual multilocus genotypes to estimate rates of recent immigration (over the last several generations) among populations is presented. The method also estimates the posterior probability distributions of individual immigrant ancestries, population allele frequencies, population inbreeding coefficients, and other parameters of potential interest. The method is implemented in a computer program that relies on Markov chain Monte Carlo techniques to carry out the estimation of posterior probabilities. The program can be used with allozyme, microsatellite, RFLP, SNP, and other kinds of genotype data. We relax several assumptions of early methods for detecting recent immigrants, using genotype data; most significantly, we allow genotype frequencies to deviate from Hardy-Weinberg equilibrium proportions within populations. The program is demonstrated by applying it to two recently published microsatellite data sets for populations of the plant species Centaurea corymbosa and the gray wolf species Canis lupus. A computer simulation study suggests that the program can provide highly accurate estimates of migration rates and individual migrant ancestries, given sufficient genetic differentiation among populations and sufficient numbers of marker loci.

Animal Migration↗

Using probabilistic and decision-theoretic methods in treatment and prognosis modeling.

Causal probabilistic networks, also called Bayesian networks, allow both qualitative knowledge about the structure of a problem and quantitative knowledge, derived from case databases, expert opinion and literature to be exploited in the construction of decision support systems for diagnosis, treatment and prognosis. This mixing of qualitative and quantitative knowledge will be illustrated, using the selection of antibiotics for a subset of patients with severe infections. The subset consists of patients where bacteria or fungi have been found in the blood. A simple pathophysiological model of infection is used to calculate a prognosis, dependent on the choice of antibiotics. A decision-theoretic approach is used to balance the therapeutic benefit of antibiotic treatment against the cost of antibiotics in the form of direct monetary cost, side effects and ecological cost. A retrospective trial on patients with bacteria or fungi in the blood stemming from the urinary tract indicates that with this approach, it may be possible to suggest balanced choices of antibiotics that not only achieve greater therapeutic benefit, but also reduce the cost of therapy.

Anti-Bacterial Agents↗

Bayesian inferences on the recent island colonization history by the bird Zosterops lateralis lateralis.

The founding of new populations by small numbers of colonists has been considered a potentially important mechanism promoting evolutionary change in island populations. Colonizing species, such as members of the avian species complex Zosterops lateralis, have been used to support this idea. A large amount of background information on recent colonization history is available for one Zosterops subspecies, Z. lateralis lateralis, providing the opportunity to reconstruct the population dynamics of its colonization sequence. We used a Bayesian approach to combine historical and demographic information available on Z. l. lateralis with genotypic data from six microsatellite loci, and a rejection algorithm to make simultaneous inferences on the demographic parameters describing the recent colonization history of this subspecies in four southwest Pacific islands. Demographic models assuming mutation-drift equilibrium or a large number of founders were better supported than models assuming founder events for three of four recently colonized island populations. Posterior distributions of demographic parameters supported (i) a large stable effective population size of several thousands individuals with point estimates around 4000-5000; (ii) a founder event of very low intensity with a large effective number of founders around 150-200 individuals for each island in three of four islands, suggesting the colonization of those islands by one flock of large size or several flocks of average size; and (iii) a founder event of higher intensity on Norfolk Island with an effective number of founders around 20 individuals, suggesting colonization by a single flock of moderate size. Our inferences on demographic parameters, especially those on the number of founders, were relatively insensitive to the precise choice of prior distributions for microsatellite mutation processes and demographic parameters, suggesting that our analysis provides a robust description of the recent colonization history of the subspecies.

Bayes Theorem↗

Association mapping of complex trait loci with context-dependent effects and unknown context variable.

A novel method for Bayesian analysis of genetic heterogeneity and multilocus association in random population samples is presented. The method is valid for quantitative and binary traits as well as for multiallelic markers. In the method, individuals are stochastically assigned into two etiological groups that can have both their own, and possibly different, subsets of trait-associated (disease-predisposing) loci or alleles. The method is favorable especially in situations when etiological models are stratified by the factors that are unknown or went unmeasured, that is, if genetic heterogeneity is due to, for example, unknown genes x environment or genes x gene interactions. Additionally, a heterogeneity structure for the phenotype does not need to follow the structure of the general population; it can have a distinct selection history. The performance of the method is illustrated with simulated example of genes x environment interaction (quantitative trait with loosely linked markers) and compared to the results of single-group analysis in the presence of missing data. Additionally, example analyses with previously analyzed cystic fibrosis and type 2 diabetes data sets (binary traits with closely linked markers) are presented. The implementation (written in WinBUGS) is freely available for research purposes from http://www.rni.helsinki.fi/ approximately mjs/.

Alleles↗

Advances to Bayesian network inference for generating causal networks from observational biological data.

MOTIVATION: Network inference algorithms are powerful computational tools for identifying putative causal interactions among variables from observational data. Bayesian network inference algorithms hold particular promise in that they can capture linear, non-linear, combinatorial, stochastic and other types of relationships among variables across multiple levels of biological organization. However, challenges remain when applying these algorithms to limited quantities of experimental data collected from biological systems. Here, we use a simulation approach to make advances in our dynamic Bayesian network (DBN) inference algorithm, especially in the context of limited quantities of biological data. RESULTS: We test a range of scoring metrics and search heuristics to find an effective algorithm configuration for evaluating our methodological advances. We also identify sampling intervals and levels of data discretization that allow the best recovery of the simulated networks. We develop a novel influence score for DBNs that attempts to estimate both the sign (activation or repression) and relative magnitude of interactions among variables. When faced with limited quantities of observational data, combining our influence score with moderate data interpolation reduces a significant portion of false positive interactions in the recovered networks. Together, our advances allow DBN inference algorithms to be more effective in recovering biological networks from experimentally collected data. AVAILABILITY: Source code and simulated data are available upon request. SUPPLEMENTARY INFORMATION: http://www.jarvislab.net/Bioinformatics/BNAdvances/

Algorithms↗

3D image segmentation of deformable objects with joint shape-intensity prior models using level sets.

We propose a novel method for 3D image segmentation, where a Bayesian formulation, based on joint prior knowledge of the object shape and the image gray levels, along with information derived from the input image, is employed. Our method is motivated by the observation that the shape of an object and the gray level variation in an image have consistent relations that provide configurations and context that aid in segmentation. We define a maximum a posteriori (MAP) estimation model using the joint prior information of the object shape and the image gray levels to realize image segmentation. We introduce a representation for the joint density function of the object and the image gray level values, and define a joint probability distribution over the variations of the object shape and the gray levels contained in a set of training images. By estimating the MAP shape of the object, we formulate the shape-intensity model in terms of level set functions as opposed to landmark points of the object shape. In addition, we evaluate the performance of the level set representation of the object shape by comparing it with the point distribution model (PDM). We found the algorithm to be robust to noise and able to handle multidimensional data, while able to avoid the need for explicit point correspondences during the training phase. Results and validation from various experiments on 2D and 3D medical images are shown.

Algorithms↗

Bayesian analysis of genetic differentiation between populations.

We introduce a Bayesian method for estimating hidden population substructure using multilocus molecular markers and geographical information provided by the sampling design. The joint posterior distribution of the substructure and allele frequencies of the respective populations is available in an analytical form when the number of populations is small, whereas an approximation based on a Markov chain Monte Carlo simulation approach can be obtained for a moderate or large number of populations. Using the joint posterior distribution, posteriors can also be derived for any evolutionary population parameters, such as the traditional fixation indices. A major advantage compared to most earlier methods is that the number of populations is treated here as an unknown parameter. What is traditionally considered as two genetically distinct populations, either recently founded or connected by considerable gene flow, is here considered as one panmictic population with a certain probability based on marker data and prior information. Analyses of previously published data on the Moroccan argan tree (Argania spinosa) and of simulated data sets suggest that our method is capable of estimating a population substructure, while not artificially enforcing a substructure when it does not exist. The software (BAPS) used for the computations is freely available from http://www.rni.helsinki.fi/~mjs.

Bayes Theorem↗

Quantifying uncertainty of predictions from cancer progression models.

MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).

Humans↗

Bayesian-regularized genetic neural networks applied to the modeling of non-peptide antagonists for the human luteinizing hormone-releasing hormone receptor.

Bayesian-regularized genetic neural networks (BRGNNs) were used to model the binding affinity (IC(50)) for 128 non-peptide antagonists for the human luteinizing hormone-releasing hormone (LHRH) receptor using 2D spatial autocorrelation vectors. As a preliminary step, a linear dependence was established by multiple linear regression (MLR) approach, selecting the relevant descriptors by genetic algorithm (GA) feature selection. The linear model showed to fit the training set (N=102) with R(2)=0.746, meanwhile BRGNN exhibited a higher value of R(2)=0.871. Beyond the improvement of training set fitting, the BRGNN model overcame the linear one by being able to describe 85% of test set (N=26) variance in comparison with 73% the MLR model. Our non-linear QSAR model illustrates the importance of an adequate distribution of atomic properties represented in topological frames and reveals the electronegativities, masses and polarizabilities as the most influencing atomic properties in the structures of the heterocycles under analysis for having an appropriate LHRH antagonistic activity. Furthermore, the ability of the non-linear selected variables for differentiating the data was evidenced when total data set was well distributed in a Kohonen self-organizing map (SOM).

Algorithms↗

Exact and approximate Bayesian estimation of net counting rates.

The stochastic fluctuations in the number of disintegrations, which had already been studied experimentally by Rutherford and other investigators at the beginning of the twentieth century, make estimation of net counting rates in the presence of background counts a challenging statistical problem. Exact and approximate Bayesian estimates of net count rates using Poisson and normal distributions for the number of counts detected during varying counting intervals are derived. The posterior densities for the net count rate are derived and plotted for uniform priors. The graphs for the exact, Poisson based, and for the approximate posterior densities of the background and net count rates, resulting from the normal approximation to the Poisson distribution, were compared. No practical differences were found when the number of observed gross counts is large. Small numerical differences in the posterior expectations and standard deviations of the counting rates appeared when the number of observed counts was small. A table showing some of these numerical differences for different background and gross counts is included. A normal approximation to the Poisson is satisfactory for the analysis of counting data when the number of observed counts is large. Some caution has to be exercised when the number of observed counts is small.

Artifacts↗

Multigrid priors for a Bayesian approach to fMRI.

We introduce multigrid priors to construct a Bayesian-inspired method to asses brain activity in functional magnetic resonance imaging (fMRI). A sequence of different scale grids is constructed over the image. Starting from the finest scale, coarse grain data variables are sequentially defined for each scale. Then we move back to finer scales, determining for each coarse scale a set of posterior probabilities. The posterior on a coarse scale is used as the prior for activity at the next finer scale. To test the method, we use a linear model with a given hemodynamic response function to construct the likelihood. We apply the method both to real and simulated data of a boxcar experiment. To measure the number of errors, we impose a decision to determine activity by setting a threshold on the posterior. Receiver operating characteristic (ROC) curves are used to study the dependence on threshold and on a few hyperparameters in the relation between specificity and sensitivity. We also study the deterioration of the results for real data, under information loss. This is done by decreasing the number of images in each period and also by decreasing the signal to noise ratio and compare the robustness to other methods.

Algorithms↗

BRCAPRO validation, sensitivity of genetic testing of BRCA1/BRCA2, and prevalence of other breast cancer susceptibility genes.

PURPOSE: To compare genetic test results for deleterious mutations of BRCA1 and BRCA2 with estimated probabilities of carrying such mutations; to assess sensitivity of genetic testing; and to assess the relevance of other susceptibility genes in familial breast and ovarian cancer. PATIENTS AND METHODS: Data analyzed were from six high-risk genetic counseling clinics and concern individuals from families for which at least one member was tested for mutations at BRCA1 and BRCA2. Predictions of genetic predisposition to breast and ovarian cancer for 301 individuals were made using BRCAPRO, a statistical model and software using Mendelian genetics and Bayesian updating. Model predictions were compared with the results of genetic testing. RESULTS: Among the test individuals, 126 were Ashkenazi Jewish, three were male subjects, 243 had breast cancer, 49 had ovarian cancer, 34 were unaffected, and 139 tested positive for BRCA1 mutations and 29 for BRCA2 mutations. BRCAPRO performed well: for the 150 probands with the smallest BRCAPRO carrier probabilities (average, 29.0%), the proportion testing positive was 32.7%; for the 151 probands with the largest carrier probabilities (average, 95.2%), 78.8% tested positive. Genetic testing sensitivity was estimated to be at least 85%, with false-negatives including mutations of susceptibility genes heretofore unknown. CONCLUSION: BRCAPRO is an accurate counseling tool for determining the probability of carrying mutations of BRCA1 and BRCA2. Genetic testing for BRCA1 and BRCA2 is highly sensitive, missing an estimated 15% of mutations. In the populations studied, breast cancer susceptibility genes other than BRCA1 and BRCA2 either do not exist, are rare, or are associated with low disease penetrance.

Adult↗

Modelling techniques and their application for monitoring in high dependency environments--learning models.

This paper reviews the use of learning models including Bayesian classifiers and artificial neural networks in monitoring and interpreting biosignals. Generally learning models applied for analysis of biosignals are "black-box' types trained on the basis of measured signals. It is illustrated that the training and application of learning models more or less follow the same sequences. The main focus is the interpretation of electrical signals from the brain (electroencephalogram (EEG) and evoked potentials (EP)). Current analysis of these signals often reveals sudden changes in the EEG or evoked potentials to be the earliest discernible signs of inadequate perfusion of the brain. They may reflect problems such as systemic arterial oxygen desaturation or hypotension arising from other body system failures during critical illness. It is suggested that these brain signals should be recorded in the critical care unit, and that they should form part of the annotated database of biosignals established during the IMPROVE project. This would allow for the development of new methods for on-line warning of impending damage to the central nervous system, such that corrective actions could be taken before permanent damage occurred.

Auscultation↗

The footprints of visual attention in the Posner cueing paradigm revealed by classification images.

In the Posner cueing paradigm, observers' performance in detecting a target is typically better in trials in which the target is present at the cued location than in trials in which the target appears at the uncued location. This effect can be explained in terms of a Bayesian observer where visual attention simply weights the information differently at the cued (attended) and uncued (unattended) locations without a change in the quality of processing at each location. Alternatively, it could also be explained in terms of visual attention changing the shape of the perceptual filter at the cued location. In this study, we use the classification image technique to compare the human perceptual filters at the cued and uncued locations in a contrast discrimination task. We did not find statistically significant differences between the shapes of the inferred perceptual filters across the two locations, nor did the observed differences account for the measured cueing effects in human observers. Instead, we found a difference in the magnitude of the classification images, supporting the idea that visual attention changes the weighting of information at the cued and uncued location, but does not change the quality of processing at each individual location.

Attention↗

Detection of mastitis in dairy cattle by use of mixture models for repeated somatic cell scores: a Bayesian approach via Gibbs sampling.

The distribution of somatic cell scores could be regarded as a mixture of at least two components depending on a cow's udder health status. A heteroscedastic two-component Bayesian normal mixture model with random effects was developed and implemented via Gibbs sampling. The model was evaluated using datasets consisting of simulated somatic cell score records. Somatic cell score was simulated as a mixture representing two alternative udder health statuses ("healthy" or "diseased"). Animals were assigned randomly to the two components according to the probability of group membership (Pm). Random effects (additive genetic and permanent environment), when included, had identical distributions across mixture components. Posterior probabilities of putative mastitis were estimated for all observations, and model adequacy was evaluated using measures of sensitivity, specificity, and posterior probability of misclassification. Fitting different residual variances in the two mixture components caused some bias in estimation of parameters. When the components were difficult to disentangle, so were their residual variances, causing bias in estimation of Pm and of location parameters of the two underlying distributions. When all variance components were identical across mixture components, the mixture model analyses returned parameter estimates essentially without bias and with a high degree of precision. Including random effects in the model increased the probability of correct classification substantially. No sizable differences in probability of correct classification were found between models in which a single cow effect (ignoring relationships) was fitted and models where this effect was split into genetic and permanent environmental components, utilizing relationship information. When genetic and permanent environmental effects were fitted, the between-replicate variance of estimates of posterior means was smaller because the model accounted for random genetic drift.

Animals↗

Computational inference of neural information flow networks.

Determining how information flows along anatomical brain pathways is a fundamental requirement for understanding how animals perceive their environments, learn, and behave. Attempts to reveal such neural information flow have been made using linear computational methods, but neural interactions are known to be nonlinear. Here, we demonstrate that a dynamic Bayesian network (DBN) inference algorithm we originally developed to infer nonlinear transcriptional regulatory networks from gene expression data collected with microarrays is also successful at inferring nonlinear neural information flow networks from electrophysiology data collected with microelectrode arrays. The inferred networks we recover from the songbird auditory pathway are correctly restricted to a subset of known anatomical paths, are consistent with timing of the system, and reveal both the importance of reciprocal feedback in auditory processing and greater information flow to higher-order auditory areas when birds hear natural as opposed to synthetic sounds. A linear method applied to the same data incorrectly produces networks with information flow to non-neural tissue and over paths known not to exist. To our knowledge, this study represents the first biologically validated demonstration of an algorithm to successfully infer neural information flow networks.

Action Potentials↗

Comparison of linear and nonlinear classification algorithms for the prediction of drug and chemical metabolism by human UDP-glucuronosyltransferase isoforms.

Partial least squares discriminant analysis (PLSDA), Bayesian regularized artificial neural network (BRANN), and support vector machine (SVM) methodologies were compared by their ability to classify substrates and nonsubstrates of 12 isoforms of human UDP-glucuronosyltransferase (UGT), an enzyme "superfamily" involved in the metabolism of drugs, nondrug xenobiotics, and endogenous compounds. Simple two-dimensional descriptors were used to capture chemical information. For each data set, 70% of the data were used for training, and the remainder were used to assess the generalization performance. In general, the SVM methodology was able to produce models with the best predictive performance, followed by BRANN and then PLSDA. However, a small number of data sets showed either equivalent or better predictability using PLSDA, which may indicate relatively linear relationships in these data sets. All SVM models showed predictive ability (>60% of test set predicted correctly) and five out of the 12 test sets showed excellent prediction (>80% prediction accuracy). These models represent the first use of pattern recognition methods to discriminate between substrates and nonsubstrates of human drug metabolizing enzymes and the first thorough assessment of three classification algorithms using multiple metabolic data sets.

Algorithms↗