Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

"Shape activity": a continuous-state HMM for moving/deforming shapes with application to abnormal activity detection.

The aim is to model "activity" performed by a group of moving and interacting objects (which can be people, cars, or different rigid components of the human body) and use the models for abnormal activity detection. Previous approaches to modeling group activity include co-occurrence statistics (individual and joint histograms) and dynamic Bayesian networks, neither of which is applicable when the number of interacting objects is large. We treat the objects as point objects (referred to as "landmarks") and propose to model their changing configuration as a moving and deforming "shape" (using Kendall's shape theory for discrete landmarks). A continuous-state hidden Markov model is defined for landmark shape dynamics in an activity. The configuration of landmarks at a given time forms the observation vector, and the corresponding shape and the scaled Euclidean motion parameters form the hidden-state vector. An abnormal activity is then defined as a change in the shape activity model, which could be slow or drastic and whose parameters are unknown. Results are shown on a real abnormal activity-detection problem involving multiple moving objects.

Algorithms↗

Finding appropriate clinical trials: evaluating encoded eligibility criteria with incomplete data.

We describe our work on creating a system that selects appropriate clinical trials by automating the evaluation of eligibility criteria. We developed a data model of eligibility for breast cancer clinical trials, upon which the criteria were encoded. Standard vocabularies are utilized to represent concepts used in the system, and retrieve their hierarchical relationships. The system incorporates Bayesian networks to handle missing patient information. Protocols are ranked by the belief that the patient is eligible for each of them. In a preliminary evaluation, we found good agreement (kappa 0.86) between the system and an independent physician in selection of protocols, but poor agreement (kappa 0.24) in protocol ranking. We conclude that our approach is feasible, and potentially useful in assisting both physicians and patients in the task of selecting appropriate trials.

Bayes Theorem↗

REVCOM: a robust Bayesian method for evolutionary rate estimation.

MOTIVATION: Evolutionary conservation estimated from a multiple sequence alignment is a powerful indicator of the functional significance of a residue and helps to predict active sites, ligand binding sites, and protein interaction interfaces. Many algorithms that calculate conservation work well, provided an accurate and balanced alignment is used. However, such a strong dependence on the alignment makes the results highly variable. We attempted to improve the conservation prediction algorithm by making it more robust and less sensitive to (1) local alignment errors, (2) overrepresentation of sequences in some branches and (3) occasional presence of unrelated sequences. RESULTS: A novel method is presented for robust constrained Bayesian estimation of evolutionary rates that avoids overfitting independent rates and satisfies the above requirements. The method is evaluated and compared with an entropy-based conservation measure on a set of 1494 protein interfaces. We demonstrated that approximately 62% of the analyzed protein interfaces are more conserved than the remaining surface at the 5% significance level. A consistent method to incorporate alignment reliability is proposed and demonstrated to reduce arbitrary variation of calculated rates upon inclusion of distantly related or unrelated sequences into the alignment.

Algorithms↗

Stochastic automata-based estimators for adaptively compressing files with nonstationary distributions.

This correspondence shows that learning automata techniques, which have been useful in developing weak estimators, can be applied to data compression applications in which the data distributions are nonstationary. The adaptive coding scheme utilizes stochastic learning-based weak estimation techniques to adaptively update the probabilities of the source symbols, and this is done without resorting to either maximum likelihood, Bayesian, or sliding-window methods. The authors have incorporated the estimator in the adaptive Fano coding scheme and in an adaptive entropy-based scheme that "resembles" the well-known arithmetic coding. The empirical results obtained for both of these adaptive methods are obtained on real-life files that possess a fair degree of nonstationarity. From these results, it can be seen that the proposed schemes compress nearly 10% more than their respective adaptive methods that use maximum-likelihood estimator-based estimates.

Algorithms↗

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem↗

A Bayesian approach to disease gene location using allelic association.

A Bayesian approach to analysing data from family-based association studies is developed. This permits direct assessment of the range of possible values of model parameters, such as the recombination frequency and allelic associations, in the light of the data. In addition, sophisticated comparisons of different models may be handled easily, even when such models are not nested. The methodology is developed in such a way as to allow separate inferences to be made about linkage and association by including theta, the recombination fraction between the marker and disease susceptibility locus under study, explicitly in the model. The method is illustrated by application to a previously published data set. The data analysis raises some interesting issues, notably with regard to the weight of evidence necessary to convince us of linkage between a candidate locus and disease.

Alleles↗

Measuring gametic disequilibrium from multilocus data.

We describe a Bayesian approach to analyzing multilocus genotype or haplotype data to assess departures from gametic (linkage) equilibrium. Our approach employs a Markov chain Monte Carlo (MCMC) algorithm to approximate the posterior probability distributions of disequilibrium parameters. The distributions are computed exactly in some simple settings. Among other advantages, posterior distributions can be presented visually, which allows the uncertainties in parameter estimates to be readily assessed. In addition, background knowledge can be incorporated, where available, to improve the precision of inferences. The method is illustrated by application to previously published datasets; implications for multilocus forensic match probabilities and for simple association-based gene mapping are also discussed.

Algorithms↗

Discovery of causal relationships in a gene-regulation pathway from a mixture of experimental and observational DNA microarray data.

This paper reports the methods and results of a computer-based search for causal relationships in the gene-regulation pathway of galactose metabolism in the yeast Saccharomyces cerevisiae. The search uses recently published data from cDNA microarray experiments. A Bayesian method was applied to learn causal networks from a mixture of observational and experimental gene-expression data. The observational data were gene-expression levels obtained from unmanipulated "wild-type" cells. The experimental data were produced by deleting ("knocking out") genes and observing the expression levels of other genes. Causal relations predicted from the analysis on 36 galactose gene pairs are reported and compared with the known galactose pathway. Additional exploratory analyses are also reported.

Animals↗

[Medical informatics as a complementary method in medical education].

The practice of the decision making at the bed side especially highlights the place to be devoted to medical informatics both at the pre- and post-graduate levels. Still in a relatively recent past, say the 50s-60s, most of the medical educational efforts were delivered when watching and then imitating the medical behaviour of an older physician. The medical educators were aware that besides the formal lessons related to selected chapters of medical textbooks, there were an obvious need for better training in the ability to make sound clinical judgements. If this ability has been considered only as an artful and intuitive process neither subjected to theoretical analysis nor to be captured in a formal quantitative model, now things have changed to such an extent that it becomes broadly shared that a science of medical decision making can be reasonably founded and this threefold: 1) Upon a formulated logic, 2) The probability theory, and 3) A value theory. The first gives the hand to artificial intelligence (AI) technics, the third to medical information data bases dealing either with patients (like in hospital information systems) or with literature like MEDLINE or electronic "cookbooks". Basically the probabilistic theory is based here upon a priori probabilities related to patients informations and data and opens the way to bayesian decision making. After this little summary it is stressed that educational informatics in medicine would appear either very central or very marginal, if not optional.

Computer-Assisted Instruction↗

[Phylogenetic analysis for H3A1 strain of all human influenza A virus].

OBJECTIVE: Influenza A virus remains an important pathogen which threatens humans. With the help of latest developed bioinformatics tools, all available human Influenza A virus H3A1 strains were explored to deeply understanding its evolution and variation rules. METHODS: All data of H3A1 sequence in NCBI Genbank and Influenza sequence database were downloaded and aligned in ClustalX with two step cluster method used to split the data and Bayesian phylogenetic tree analysis method applied to precisely construct phylogenetic tree for each clusters. RESULTS: Tree topology indicated that H3 strains evolved along a single evolution trunk and tree pattern and model parameter showed obvious variety tendency with time period. However, no geographic distribution features were found for key variation strains and big branch in trees. CONCLUSION: The evolution of human H3 strains were mainly driven by the interaction of human immune barriers and antigenic drift of virus. Since the influenza subtype had already been spread in human population, south China should not be considered as the originated areas of new strains, hence it should be treated as equally as other places in the world.

Antigens↗

Experimental design of time series data for learning from dynamic Bayesian networks.

Bayesian networks (BNs) and dynamic Bayesian networks (DBNs) are becoming more widely used as a way to learn various types of networks, including cellular signaling networks, from high-throughput data. Due to the high cost of performing experiments, we are interested in developing an experimental design for time series data generation. Specifically, we are interested in determining properties of time series data that make them more efficient for DBN modeling. We present a theoretical analysis on the ability of DBNs without hidden variables to learn from proteomic time series data. The analysis reveals, among other lessons, that under a reasonable set of assumptions a fixed budget is better spent on collecting many short time series data than on a few long time series data.

Algorithms↗

A unified framework for subspace face recognition.

PCA, LDA, and Bayesian analysis are the three most representative subspace face recognition approaches. In this paper, we show that they can be unified under the same framework. We first model face difference with three components: intrinsic difference, transformation difference, and noise. A unified framework is then constructed by using this face difference model and a detailed subspace analysis on the three components. We explain the inherent relationship among different subspace methods and their unique contributions to the extraction of discriminating information from the face difference. Based on the framework, a unified subspace analysis method is developed using PCA, Bayes, and LDA as three steps. A 3D parameter space is constructed using the three subspace dimensions as axes. Searching through this parameter space, we achieve better recognition performance than standard subspace methods.

Algorithms↗

Markov random field models for directional field and singularity extraction in fingerprint images.

A Bayesian formulation is proposed for reliable and robust extraction of the directional field in fingerprint images using a class of spatially smooth priors. The spatial smoothness allows for robust directional field estimation in the presence of moderate noise levels. Parametric template models are suggested as candidate singularity models for singularity detection. The parametric models enable joint extraction of the directional field and the singularities in fingerprint impressions by dynamic updating of feature information. This allows for the detection of singularities that may have previously been missed, as well as better aligning the directional field around detected singularities. A criteria is presented for selecting an optimal block size to reduce the number of spurious singularity detections. The best rates of spurious detection and missed singularities given by the algorithm are 4.9% and 7.1%, respectively, based on the NIST 4 database.

Algorithms↗

Source continuity and boundary discontinuity considerations in Bayesian image processing.

This paper extends the Bayesian image processing (BIP) formalism by considering the effect of simple source continuity and boundary discontinuity and a priori information in estimating an optimal source distribution from observed data. The a priori source information is formulated in terms of probability density functions of source element strengths and spatial correlations. The estimation is carried out iteratively by a BIP algorithm derived by applying the expectation maximization technique to the a priori source probability density functions and assuming the data obey Poisson statistics. The suppression of boundary oscillations and enhancement of overall image are demonstrated for computer generated ideal and Poisson randomized data.

Algorithms↗

Predicting phenytoin dosages using Bayesian feedback: a comparison with other methods.

A Bayesian feedback technique for predicting phenytoin dosage was compared to other dosing methods. Sixty-nine cases were selected on the basis of apparent reliability from 103 medical charts of epileptic patients with multiple phenytoin levels on different dosage regimens. Two published nomograms and a graphical, or computational, technique were compared to the Bayesian technique. Each method was assessed for absolute predictability using measures of bias and precision, i.e., mean percent error and root mean squared percent error, respectively. For a single previous data pair, the Bayesian method was similar to a published nomogram with regard to bias and precision. For multiple data pairs, the graphical or simultaneous equation technique tended to be less biased, but the Bayesian method had better precision. However, none of these differences was statistically significant (p greater than 0.05). The Bayesian method yielded the lowest percentage of predicted doses that exceeded 110% of the actual dose. The Bayesian method conveniently provides a single method applicable to the use of either single or multiple concentration-dosage data pairs and results in fewer extreme dosing errors.

Adult↗

Bayesian estimation, simulation and uncertainty analysis: the cost-effectiveness of ganciclovir prophylaxis in liver transplantation.

This paper demonstrates the usefulness of combining simulation with Bayesian estimation methods in analysis of cost-effectiveness data collected alongside a clinical trial. Specifically, we use Markov Chain Monte Carlo (MCMC) to estimate a system of generalized linear models relating costs and outcomes to a disease process affected by treatment under alternative therapies. The MCMC draws are used as parameters in simulations which yield inference about the relative cost-effectiveness of the novel therapy under a variety of scenarios. Total parametric uncertainty is assessed directly by examining the joint distribution of simulated average incremental cost and effectiveness. The approach allows flexibility in assessing treatment in various counterfactual premises and quantifies the global effect of parametric uncertainty on a decision-maker's confidence in adopting one therapy over the other.

Antiviral Agents↗

Bayesian segmental models with multiple sequence alignment profiles for protein secondary structure and contact map prediction.

In this paper, we develop a segmental semi-Markov model (SSMM) for protein secondary structure prediction which incorporates multiple sequence alignment profiles with the purpose of improving the predictive performance. The segmental model is a generalization of the hidden Markov model where a hidden state generates segments of various length and secondary structure type. A novel parameterized model is proposed for the likelihood function that explicitly represents multiple sequence alignment profiles to capture the segmental conformation. Numerical results on benchmark data sets show that incorporating the profiles results in substantial improvements and the generalization performance is promising. By incorporating the information from long range interactions in beta-sheets, this model is also capable of carrying out inference on contact maps. This is an important advantage of probabilistic generative models over the traditional discriminative approach to protein secondary structure prediction. The Web server of our algorithm and supplementary materials are available at http://public.kgi.edu/-wild/bsm.html.

Algorithms↗

BACUS: A Bayesian protocol for the identification of protein NOESY spectra via unassigned spin systems.

NMR frequency assignments are usually considered a prerequisite for the analysis of NOESY spectra, in turn required for the calculation of biomolecular structures. In contrast, as we propose here, relatively high numbers of unambiguous NOE identities can be consistently achieved in an automated manner by relying only on grouping resonances into connected spin systems. To achieve this goal, we have developed for proteins two protocols, SPI and BACUS, based on Bayesian inference. SPI (Grishaev and Llinás, 2002c) produces a list of the (1)H resonance frequencies from homo- and hetero-nuclear multidimensional spectra, grouped into effective spin systems. BACUS automatically establishes probabilistic identities of NOESY cross-peaks in terms of the chemical shifts provided by SPI. BACUS requires neither assignment of resonances nor an initial structural model. It successfully copes with chemical shift overlap and does so without cycling through 3D structure calculations. The method exploits the self-consistency of the NOESY graph by taking advantage of a network of J- as well as NOE-connected "reporter" protons sorted via SPI. BACUS was validated by tests on experimental NOESY data recorded for the col 2 and kringle 2 domains.

Bayes Theorem↗