Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Bayesian inference applied to macromolecular structure determination.

The determination of macromolecular structures from experimental data is an ill-posed inverse problem. Nevertheless, conventional techniques to structure determination attempt an inversion of the data by minimization of a target function. This approach leads to problems if the data are sparse, noisy, heterogeneous, or difficult to describe theoretically. We propose here to view biomolecular structure determination as an inference rather than an inversion problem. Probability theory then offers a consistent formalism to solve any structure determination problem: We use Bayes' theorem to derive a probability distribution for the atomic coordinates and all additional unknowns. This distribution represents the complete information contained in the data and can be analyzed numerically by Markov chain Monte Carlo sampling techniques. We apply our method to data obtained from a nuclear magnetic resonance experiment and discuss the estimation of theory parameters.

Bayes Theorem↗

Bayesian estimation of fold-changes in the analysis of gene expression: the PFOLD algorithm.

A general and detailed noise model for the DNA microarray measurement of gene expression is presented and used to derive a Bayesian estimation scheme for expression ratios, implemented in a program called PFOLD, which provides not only an estimate of the fold-change in gene expression, but also confidence limits for the change and a P-value quantifying the significance of the change. Although the focus is on oligonucleotide microarray technologies, the scheme can also be applied to cDNA based technologies if parameters for the noise model are provided. The model unifies estimation for all signals in that it provides a seamless transition from very low to very high signal-to-noise ratios, an essential feature for current microarray technologies for which the median signal-to-noise ratios are always moderate. The dual use, as decision statistics in a two-dimensional space, of the P-value and the fold-change is shown to be effective in the ubiquitous problem of detecting changing genes against a background of unchanging genes, leading to markedly higher sensitivities, at equal selectivity, than detection and selection based on the fold-change alone, a current practice until now.

Algorithms↗

Bayesian applications of belief networks and multilayer perceptrons for ovarian tumor classification with rejection.

Incorporating prior knowledge into black-box classifiers is still much of an open problem. We propose a hybrid Bayesian methodology that consists in encoding prior knowledge in the form of a (Bayesian) belief network and then using this knowledge to estimate an informative prior for a black-box model (e.g. a multilayer perceptron). Two technical approaches are proposed for the transformation of the belief network into an informative prior. The first one consists in generating samples according to the most probable parameterization of the Bayesian belief network and using them as virtual data together with the real data in the Bayesian learning of a multilayer perceptron. The second approach consists in transforming probability distributions over belief network parameters into distributions over multilayer perceptron parameters. The essential attribute of the hybrid methodology is that it combines prior knowledge and statistical data efficiently when prior knowledge is available and the sample is of small or medium size. Additionally, we describe how the Bayesian approach can provide uncertainty information about the predictions (e.g. for classification with rejection). We demonstrate these techniques on the medical task of predicting the malignancy of ovarian masses and summarize the practical advantages of the Bayesian approach. We compare the learning curves for the hybrid methodology with those of several belief networks and multilayer perceptrons. Furthermore, we report the performance of Bayesian belief networks when they are allowed to exclude hard cases based on various measures of prediction uncertainty.

Bayes Theorem↗

A Bayesian method for finding regulatory segments in DNA.

A goal of the human genome project is to determine the entire sequence of DNA (3 x 10(9) base pairs) found in chromosomes. The massive amounts of data produced by this project require interpretation. A Bayesian model is developed for locating regulatory regions in a DNA sequence. Regulatory regions are areas of DNA to which specific proteins bind and control whether or not a gene is transcribed to produce templates for protein synthesis. Each human cell contains the same DNA sequence. Thus the particular function of different cells is determined by the genes that are transcribed in that cell. A Hidden Markov chain is used to model whether a small interval of the DNA is in a regulatory region or not. This can be regarded as a changepoint problem where the changepoints are the start of a regulatory or nonregulatory region. The data consists of protein-binding elements, which are short subsequences, or "words," in the DNA sequence. Although these words can occur anywhere in the sequence, a larger number are expected in regulatory regions. Therefore, regulatory regions are detected by locating clusters of words. For a particular DNA sequence, the model automatically selects those words that best predict regions of interest. Markov chain Monte Carlo methods are used to explore the posterior distribution of the Hidden Markov chain. The model is tested by means of simulations, and applied to several DNA sequences.

Bayes Theorem↗

The cluster variation method for efficient linkage analysis on extended pedigrees.

BACKGROUND: Computing exact multipoint LOD scores for extended pedigrees rapidly becomes infeasible as the number of markers and untyped individuals increase. When markers are excluded from the computation, significant power may be lost. Therefore accurate approximate methods which take into account all markers are desirable. METHODS: We present a novel method for efficient estimation of LOD scores on extended pedigrees. Our approach is based on the Cluster Variation Method, which deterministically estimates likelihoods by performing exact computations on tractable subsets of variables (clusters) of a Bayesian network. First a distribution over inheritances on the marker loci is approximated with the Cluster Variation Method. Then this distribution is used to estimate the LOD score for each location of the trait locus. RESULTS: First we demonstrate that significant power may be lost if markers are ignored in the multi-point analysis. On a set of pedigrees where exact computation is possible we compare the estimates of the LOD scores obtained with our method to the exact LOD scores. Secondly, we compare our method to a state of the art MCMC sampler. When both methods are given equal computation time, our method is more efficient. Finally, we show that CVM scales to large problem instances. CONCLUSION: We conclude that the Cluster Variation Method is as accurate as MCMC and generally is more efficient. Our method is a promising alternative to approaches based on MCMC sampling.

Alleles↗

Automated design of robust discriminant analysis classifier for foot pressure lesions using kinematic data.

In the recent years, the use of motion tracking systems for acquisition of functional biomechanical gait data, has received increasing interest due to the richness and accuracy of the measured kinematic information. However, costs frequently restrict the number of subjects employed, and this makes the dimensionality of the collected data far higher than the available samples. This paper applies discriminant analysis algorithms to the classification of patients with different types of foot lesions, in order to establish an association between foot motion and lesion formation. With primary attention to small sample size situations, we compare different types of Bayesian classifiers and evaluate their performance with various dimensionality reduction techniques for feature extraction, as well as search methods for selection of raw kinematic variables. Finally, we propose a novel integrated method which fine-tunes the classifier parameters and selects the most relevant kinematic variables simultaneously. Performance comparisons are using robust resampling techniques such as Bootstrap 632+ and k-fold cross-validation. Results from experimentations with lesion subjects suffering from pathological plantar hyperkeratosis, show that the proposed method can lead to approximately 96% correct classification rates with less than 10% of the original features.

Adult↗

Refining protein subcellular localization.

The study of protein subcellular localization is important to elucidate protein function. Even in well-studied organisms such as yeast, experimental methods have not been able to provide a full coverage of localization. The development of bioinformatic predictors of localization can bridge this gap. We have created a Bayesian network predictor called PSLT2 that considers diverse protein characteristics, including the combinatorial presence of InterPro motifs and protein interaction data. We compared the localization predictions of PSLT2 to high-throughput experimental localization datasets. Disagreements between these methods generally involve proteins that transit through or reside in the secretory pathway. We used our multi-compartmental predictions to refine the localization annotations of yeast proteins primarily by distinguishing between soluble lumenal proteins and soluble proteins peripherally associated with organelles. To our knowledge, this is the first tool to provide this functionality. We used these sub-compartmental predictions to characterize cellular processes on an organellar scale. The integration of diverse protein characteristics and protein interaction data in an appropriate setting can lead to high-quality detailed localization annotations for whole proteomes. This type of resource is instrumental in developing models of whole organelles that provide insight into the extent of interaction and communication between organelles and help define organellar functionality.

Amino Acid Motifs↗

Bayesian two-compartment and classic single-compartment minimal models: comparison on insulin modified IVGTT and effect of experiment reduction.

Models describing plasma glucose and insulin concentration of an intravenous glucose tolerance test (IVGTT) allow a noninvasive cost-effective approach to estimate important indexes characterizing the efficiency of glucose-insulin control system, i.e., glucose effectiveness (S(G)) and insulin sensitivity (S(I)). To overcome some limitations of the classic single compartment minimal model (1CMM) of glucose kinetics , a two-compartment Bayesian minimal model (2CBMM) has been recently proposed for the standard IVGTT. This study aims to assess 2CBMM ability to describe the insulin-modified IVGTT (IM-IVGTT) which is the protocol of choice since it allows to study insulinopenic states. Both a full-length IM-IVGTT (240 min) as well as a reduced version (90 min) of it are studied. Results of the maximum a posteriori identification of IM-IVGTT (240 min) in 13 normals agree with those of standard IVGTT, i.e., a 42% decrease (P < 0.002) of S(G) and a 13% increase (P < 0.006) of S(I) with respect to ICMM. When identified from IM-IVGTT (90 min), 2CBMM not only provides S(G) and S(I) estimates 46% lower (P < 0.002) and 41% higher (P < 0.002) than 1CMM ones respectively, but also seems to overcome some limitations of the 240 min-based identification that probably arise because the minimal model is unable to properly account for the hyperglycemic hormonal response taking place in the second half of IM-IVGTT.

Bayes Theorem↗

Gene network inference from incomplete expression data: transcriptional control of hematopoietic commitment.

MOTIVATION: The topology and function of gene regulation networks are commonly inferred from time series of gene expression levels in cell populations. This strategy is usually invalid if the gene expression in different cells of the population is not synchronous. A promising, though technically more demanding alternative is therefore to measure the gene expression levels in single cells individually. The inference of a gene regulation network requires knowledge of the gene expression levels at successive time points, at least before and after a network transition. However, owing to experimental limitations a complete determination of the precursor state is not possible. RESULTS: We investigate a strategy for the inference of gene regulatory networks from incomplete expression data based on dynamic Bayesian networks. This permits prediction of the number of experiments necessary for network inference depending on parameters including noise in the data, prior knowledge and limited attainability of initial states. Our strategy combines a gradual 'Partial Learning' approach based solely on true experimental observations for the network topology with expectation maximization for the network parameters. We illustrate our strategy by extensive computer simulations in a high-dimensional parameter space in a simulated single-cell-based example of hematopoietic stem cell commitment and in random networks of different sizes. We find that the feasibility of network inferences increases significantly with the experimental ability to force the system into different initial network states, with prior knowledge and with noise reduction. AVAILABILITY: Source code is available under: www.izbi.uni-leipzig.de/services/NetwPartLearn.html SUPPLEMENTARY INFORMATION: Supplementary Data are available at Bioinformatics online.

Algorithms↗

An image-processing system for qualitative and quantitative volumetric analysis of brain images.

In this work, we developed, implemented, and validated an image-processing system for qualitative and quantitative volumetric analysis of brain images. This system allows the visualization and quantitation of global and regional brain volumes. Global volumes were obtained via an automated adaptive Bayesian segmentation technique that labels the brain into white matter, gray matter, and cerebrospinal fluid. Absolute volumetric errors for these compartments ranged between 1 and 3% as indicated by phantom studies. Quantitation of regional brain volumes was performed through normalization and tessellation of segmented brain images into the Talairach space with a 3D elastic warping model. Retest reliability of regional volumes measured in Talairach space indicated errors of < 1.5% for the frontal, parietal, temporal, and occipital brain regions. Additional regional analysis was performed with an automated hybrid method combining a region-of-interest approach and voxel-based analysis, named Regional Analysis of Volumes Examined in Normalized Space (RAVENS). RAVENS analysis for several subcortical structures showed good agreement with operator-defined volumes. This system has sufficient accuracy for longitudinal imaging data and is currently being used in the analysis of neuroimaging data of the Baltimore Longitudinal Study of Aging.

Aged↗

Detecting the number of clusters of individuals using the software STRUCTURE: a simulation study.

The identification of genetically homogeneous groups of individuals is a long standing issue in population genetics. A recent Bayesian algorithm implemented in the software STRUCTURE allows the identification of such groups. However, the ability of this algorithm to detect the true number of clusters (K) in a sample of individuals when patterns of dispersal among populations are not homogeneous has not been tested. The goal of this study is to carry out such tests, using various dispersal scenarios from data generated with an individual-based model. We found that in most cases the estimated 'log probability of data' does not provide a correct estimation of the number of clusters, K. However, using an ad hoc statistic DeltaK based on the rate of change in the log probability of data between successive K values, we found that STRUCTURE accurately detects the uppermost hierarchical level of structure for the scenarios we tested. As might be expected, the results are sensitive to the type of genetic marker used (AFLP vs. microsatellite), the number of loci scored, the number of populations sampled, and the number of individuals typed in each sample.

Bayes Theorem↗

A Bayesian approach to inferring population structure from dominant markers.

Molecular markers derived from polymerase chain reaction (PCR) amplification of genomic DNA are an important part of the toolkit of evolutionary geneticists. Random amplified polymorphic DNA markers (RAPDs), amplified fragment length polymorphisms (AFLPs) and intersimple sequence repeat (ISSR) polymorphisms allow analysis of species for which previous DNA sequence information is lacking, but dominance makes it impossible to apply standard techniques to calculate F-statistics. We describe a Bayesian method that allows direct estimates of FST from dominant markers. In contrast to existing alternatives, we do not assume previous knowledge of the degree of within-population inbreeding. In particular, we do not assume that genotypes within populations are in Hardy-Weinberg proportions. Our estimate of FST incorporates uncertainty about the magnitude of within-population inbreeding. Simulations show that samples from even a relatively small number of loci and populations produce reliable estimates of FST. Moreover, some information about the degree of within-population inbreeding (FIS) is available from data sets with a large number of loci and populations. We illustrate the method with a reanalysis of RAPD data from 14 populations of a North American orchid, Platanthera leucophaea.

Bayes Theorem↗

Hemodynamic segmentation of MR brain perfusion images using independent component analysis, thresholding, and Bayesian estimation.

Dynamic-susceptibility-contrast MR perfusion imaging is a widely used imaging tool for in vivo study of cerebral blood perfusion. However, visualization of different hemodynamic compartments is less investigated. In this work, independent component analysis, thresholding, and Bayesian estimation were used to concurrently segment different tissues, i.e., artery, gray matter, white matter, vein and sinus, choroid plexus, and cerebral spinal fluid, with corresponding signal-time curves on perfusion images of five normal volunteers. Based on the spatiotemporal hemodynamics, sequential passages and microcirculation of contrast-agent particles in these tissues were decomposed and analyzed. Late and multiphasic perfusion, indicating the presence of contrast agents, was observed in the choroid plexus and the cerebral spinal fluid. An arterial input function was modeled using the concentration-time curve of the arterial area on the same slice, rather than remote slices, for the deconvolution calculation of relative cerebral blood flow.

Adolescent↗

EscaPRRS-ORF5: a structure-aware evolutionary framework for prioritizing immune escape-prone variants in porcine reproductive and respiratory syndrome virus.

MOTIVATION: Porcine Reproductive and Respiratory Syndrome Virus (PRRSV) is a rapidly evolving RNA virus causing significant economic losses, posing a formidable challenge to vaccine efficacy due to its high mutational variability and immune escape. As the viral mutants evolve, their ability to sustain in population is driven by a range of host biology factors such as receptor binding, fusion, and uncoating. Existing tools that predict viral fitness and escape propensities rely heavily on extensive, up-to-date sequence data and lack integration of biochemical host interactions, limiting mechanistic understanding of the mutational landscape. We introduce Esca, a sequence-only toolchain framework that identifies immune escape-prone residues by exhaustively scanning each residue position for all amino acid substitutions using a Bayesian Variational Autoencoder (VAE) trained on protein language model embeddings. We demonstrate Esca on the GP5(ORF5) glycoprotein of PRRSV (EscaPRRS-ORF5) by training on ESM-2 embeddings of 32&#x2006;146 GP5 sequences (2015-2022) spanning 140 sub-lineages. RESULTS: Despite being trained only on GP5 sequence data, EscaPRRS-ORF5 recovered 85.7% of the surface-exposed receptor binding interfaces as escape-prone regions. We use a mutation-sensitive fitness scoring scheme that goes beyond Hamming distances, to predict antibody escape tendencies, supporting surveillance of (re) emerging PRRSV variants. We do not claim that ORF5 alone captures PRRSV evolution or serves as a surveillance endpoint; rather, Esca offers a scalable path toward whole-genome, structure-aware surveillance. AVAILABILITY AND IMPLEMENTATION: EscaPRRS-ORF5 is freely available at https://doi.org/10.6084/m9.figshare.32661033 with an interactive Colab notebook at https://colab.research.google.com/drive/1TEgzAhPwvNAZ01VXeJbIFibfri2jnDA5? usp=sharing.

Porcine respiratory and reproductive syndrome viru↗

A Bayesian framework for extracting human gait using strong prior knowledge.

Extracting full-body motion of walking people from monocular video sequences in complex, real-world environments is an important and difficult problem, going beyond simple tracking, whose satisfactory solution demands an appropriate balance between use of prior knowledge and learning from data. We propose a consistent Bayesian framework for introducing strong prior knowledge into a system for extracting human gait. In this work, the strong prior is built from a simple articulated model having both time-invariant (static) and time-variant (dynamic) parameters. The model is easily modified to cater to situations such as walkers wearing clothing that obscures the limbs. The statistics of the parameters are learned from high-quality (indoor laboratory) data and the Bayesian framework then allows us to "bootstrap" to accurate gait extraction on the noisy images typical of cluttered, outdoor scenes. To achieve automatic fitting, we use a hidden Markov model to detect the phases of images in a walking cycle. We demonstrate our approach on silhouettes extracted from fronto-parallel ("sideways on") sequences of walkers under both high-quality indoor and noisy outdoor conditions. As well as high-quality data with synthetic noise and occlusions added, we also test walkers with rucksacks, skirts, and trench coats. Results are quantified in terms of chamfer distance and average pixel error between automatically extracted body points and corresponding hand-labeled points. No one part of the system is novel in itself, but the overall framework makes it feasible to extract gait from very much poorer quality image sequences than hitherto. This is confirmed by comparing person identification by gait using our method and a well-established baseline recognition algorithm.

Algorithms↗

Compensation of log-compressed images for 3-D ultrasound.

In this study, a Bayesian approach was used for 3-D reconstruction in the presence of multiplicative noise and nonlinear compression of the ultrasound (US) data. Ultrasound images are often considered as being corrupted by multiplicative noise (speckle). Several statistical models have been developed to represent the US data. However, commercial US equipment performs a nonlinear image compression that reduces the dynamic range of the US signal for visualization purposes. This operation changes the distribution of the image pixels, preventing a straightforward application of the models. In this paper, the nonlinear compression is explicitly modeled and considered in the reconstruction process, where the speckle noise present in the radio frequency (RF) US data is modeled with a Rayleigh distribution. The results obtained by considering the compression of the US data are then compared with those obtained assuming no compression. It is shown that the estimation performed using the nonlinear log-compression model leads to better results than those obtained with the Rayleigh reconstruction method. The proposed algorithm is tested with synthetic and real data and the results are discussed. The results have shown an improvement in the reconstruction results when the compression operation is included in the image formation model, leading to sharper images with enhanced anatomical details.

Algorithms↗

Automated semantic analysis of changes in image sequences of neurons in culture.

Quantitative studies of dynamic behaviors of live neurons are currently limited by the slowness, subjectivity, and tedium of manual analysis of changes in time-lapse image sequences. Challenges to automation include the complexity of the changes of interest, the presence of obfuscating and uninteresting changes due to illumination variations and other imaging artifacts, and the sheer volume of recorded data. This paper describes a highly automated approach that not only detects the interesting changes selectively, but also generates quantitative analyses at multiple levels of detail. Detailed quantitative neuronal morphometry is generated for each frame. Frame-to-frame neuronal changes are measured and labeled as growth, shrinkage, merging, or splitting, as would be done by a human expert. Finally, events unfolding over longer durations, such as apoptosis and axonal specification, are automatically inferred from the short-term changes. The proposed method is based on a Bayesian model selection criterion that leverages a set of short-term neurite change models and takes into account additional evidence provided by an illumination-insensitive change mask. An automated neuron tracing algorithm is used to identify the objects of interest in each frame. A novel curve distance measure and weighted bipartite graph matching are used to compare and associate neurites in successive frames. A separate set of multi-image change models drives the identification of longer term events. The method achieved frame-to-frame change labeling accuracies ranging from 85% to 100% when tested on 8 representative recordings performed under varied imaging and culturing conditions, and successfully detected all higher order events of interest. Two sequences were used for training the models and tuning their parameters; the learned parameter settings can be applied to hundreds of similar image sequences, provided imaging and culturing conditions are similar to the training set. The proposed approach is a substantial innovation over manual annotation and change analysis, accomplishing in minutes what it would take an expert hours to complete.

Algorithms↗

Bayesian modeling of dynamic scenes for object detection.

Accurate detection of moving objects is an important precursor to stable tracking or recognition. In this paper, we present an object detection scheme that has three innovations over existing approaches. First, the model of the intensities of image pixels as independent random variables is challenged and it is asserted that useful correlation exists in intensities of spatially proximal pixels. This correlation is exploited to sustain high levels of detection accuracy in the presence of dynamic backgrounds. By using a nonparametric density estimation method over a joint domain-range representation of image pixels, multimodal spatial uncertainties and complex dependencies between the domain (location) and range (color) are directly modeled. We propose a model of the background as a single probability density. Second, temporal persistence is proposed as a detection criterion. Unlike previous approaches to object detection which detect objects by building adaptive models of the background, the foreground is modeled to augment the detection of objects (without explicit tracking) since objects detected in the preceding frame contain substantial evidence for detection in the current frame. Finally, the background and foreground models are used competitively in a MAP-MRF decision framework, stressing spatial context as a condition of detecting interesting objects and the posterior function is maximized efficiently by finding the minimum cut of a capacitated graph. Experimental validation of the proposed method is performed and presented on a diverse set of dynamic scenes.

Algorithms↗