Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Bayesian modeling of time-varying and waning exposure effects.

In epidemiologic studies, there is often interest in assessing the association between exposure history and disease incidence. For many diseases, incidence may depend not only on cumulative exposure, but also on the ages at which exposure occurred. This article proposes a flexible Bayesian approach for modeling age-varying and waning exposure effects. The Cox model is generalized to allow the hazard of disease to depend on an integral, across the exposed ages, of a piecewise polynomial function of age, multiplied by an exponential decay term. Linearity properties of the model facilitate posterior computation via a Gibbs sampler, which generalizes previous algorithms for Cox regression with time-dependent covariates. The approach is illustrated by an application to the study of protective effects of breastfeeding on incidence of childhood asthma.

Adult↗

Lifting a veil on diversity: a Bayesian approach to fitting relative-abundance models.

Bayesian methods incorporate prior knowledge into a statistical analysis. This prior knowledge is usually restricted to assumptions regarding the form of probability distributions of the parameters of interest, leaving their values to be determined mainly through the data. Here we show how a Bayesian approach can be applied to the problem of drawing inference regarding species abundance distributions and comparing diversity indices between sites. The classic log series and the lognormal models of relative- abundance distribution are apparently quite different in form. The first is a sampling distribution while the other is a model of abundance of the underlying population. Bayesian methods help unite these two models in a common framework. Markov chain Monte Carlo simulation can be used to fit both distributions as small hierarchical models with shared common assumptions. Sampling error can be assumed to follow a Poisson distribution. Species not found in a sample, but suspected to be present in the region or community of interest, can be given zero abundance. This not only simplifies the process of model fitting, but also provides a convenient way of calculating confidence intervals for diversity indices. The method is especially useful when a comparison of species diversity between sites with different sample sizes is the key motivation behind the research. We illustrate the potential of the approach using data on fruit-feeding butterflies in southern Mexico. We conclude that, once all assumptions have been made transparent, a single data set may provide support for the belief that diversity is negatively affected by anthropogenic forest disturbance. Bayesian methods help to apply theory regarding the distribution of abundance in ecological communities to applied conservation.

Animals↗

Bayesian inference on biopolymer models.

MOTIVATION: Most existing bioinformatics methods are limited to making point estimates of one variable, e.g. the optimal alignment, with fixed input values for all other variables, e.g. gap penalties and scoring matrices. While the requirement to specify parameters remains one of the more vexing issues in bioinformatics, it is a reflection of a larger issue: the need to broaden the view on statistical inference in bioinformatics. RESULTS: The assignment of probabilities for all possible values of all unknown variables in a problem in the form of a posterior distribution is the goal of Bayesian inference. Here we show how this goal can be achieved for most bioinformatics methods that use dynamic programming. Specifically, a tutorial style description of a Bayesian inference procedure for segmentation of a sequence based on the heterogeneity in its composition is given. In addition, full Bayesian inference algorithms for sequence alignment are described. AVAILABILITY: Software and a set of transparencies for a tutorial describing these ideas are available at http://www.wadsworth.org/res&res/bioinfo/

Bayes Theorem↗

Decoding spike trains instant by instant using order statistics and the mixture-of-Poissons model.

In the brain, spike trains are generated in time and presumably also interpreted as they unfold in time. Recent work (Oram et al., 1999; Baker and Lemon, 2000) suggests that in several areas of the monkey brain, individual spike times carry information because they reflect an underlying rate variation. Constructing a model based on this stochastic structure allows us to apply order statistics to decode spike trains instant by instant as spikes arrive or do not. Order statistics are time-consuming to compute in the general case. We demonstrate that data from neurons in primary visual cortex are well fit by a mixture of Poisson processes; in this special case, our computations are substantially faster. In these data, spike timing contributed information beyond that available from the spike count throughout the trial. At the end of the trial, a decoder based on the mixture-of-Poissons model correctly decoded about three times as many trials as expected by chance, compared with approximately twice as many as expected by chance using the spike count only. If our model perfectly described the spike trains, and enough data were available to estimate model parameters, then our Bayesian decoder would be optimal. For four-fifths of the sets of stimulus-elicited responses, the observed spike trains were consistent with the mixture-of-Poissons model. Most of the error in estimating stimulus probabilities is attributable to not having enough data to specify the parameters of the model rather than to misspecification of the model itself.

Action Potentials↗

A quantization method based on threshold optimization for microarray short time series.

BACKGROUND: Reconstructing regulatory networks from gene expression profiles is a challenging problem of functional genomics. In microarray studies the number of samples is often very limited compared to the number of genes, thus the use of discrete data may help reducing the probability of finding random associations between genes. RESULTS: A quantization method, based on a model of the experimental error and on a significance level able to compromise between false positive and false negative classifications, is presented, which can be used as a preliminary step in discrete reverse engineering methods. The method is tested on continuous synthetic data with two discrete reverse engineering methods: Reveal and Dynamic Bayesian Networks. CONCLUSION: The quantization method, evaluated in comparison with two standard methods, 5% threshold based on experimental error and rank sorting, improves the ability of Reveal and Dynamic Bayesian Networks to identify relations among genes.

Algorithms↗

Bayesian estimation of p-aminohippurate clearance by a limited sampling strategy.

This study describes a methodology to calculated p-aminohippurate (PAH) clearance (CL) and volume of distribution (V) with both the population parameters and one or two samples taken during the disposition and the elimination phase after a single intravenous infusion. The computer program P-PHARM was used, and a log-normal distribution and a heteroscedastic residual error distribution were assumed. Ninety-six patients with and without renal insufficiency were available for analysis, and a two-compartment model was used for data modeling. Population parameters were evaluated for 70 patients (mean number of observed concentration per individual, 6) by a three-step approach. In step 1, the computer program was used to estimate the average pharmacokinetic parameters without taking into account the demographic and/or biological factors. In step 2, the relationship between the posterior individual estimates and the covariables was investigated with multiple linear stepwise algorithm. In step 3, the population parameters were re-estimated considering the relationship with the covariables. From the regression performed in step 2, the following covariables were included: serum creatinine, body surface area, and body weight. The population averages of CL and V were 30.7 +/- 2.36 L/h and 10.6 +/- 1.29 L, respectively. To evaluate the predictive performance of the population parameters, the remaining 26 patients were used. The population parameters combined with one or two individual PAH plasma concentrations led to a bayesian estimation of individual CL and V. This estimation was compared with the classical procedure of parameter estimation (individual fitting from multiple blood samples).(ABSTRACT TRUNCATED AT 250 WORDS)

Adult↗

Particle filters, a quasi-Monte-Carlo-solution for segmentation of coronaries.

In this paper we propose a Particle Filter-based approach for the segmentation of coronary arteries. To this end, successive planes of the vessel are modeled as unknown states of a sequential process. Such states consist of the orientation, position, shape model and appearance (in statistical terms) of the vessel that are recovered in an incremental fashion, using a sequential Bayesian filter (Particle Filter). In order to account for bifurcations and branchings, we consider a Monte Carlo sampling rule that propagates in parallel multiple hypotheses. Promising results on the segmentation of coronary arteries demonstrate the potential of the proposed approach.

Algorithms↗

Calculating the evolutionary rates of different genes: a fast, accurate estimator with applications to maximum likelihood phylogenetic analysis.

In phylogenetic analyses with combined multigene or multiprotein data sets, accounting for differing evolutionary dynamics at different loci is essential for accurate tree prediction. Existing maximum likelihood (ML) and Bayesian approaches are computationally intensive. We present an alternative approach that is orders of magnitude faster. The method, Distance Rates (DistR), estimates rates based upon distances derived from gene/protein sequence data. Simulation studies indicate that this technique is accurate compared with other methods and robust to missing sequence data. The DistR method was applied to a fungal mitochondrial data set, and the rate estimates compared well to those obtained using existing ML and Bayesian approaches. Inclusion of the protein rates estimated from the DistR method into the ML calculation of trees as a branch length multiplier resulted in a significantly improved fit as measured by the Akaike Information Criterion (AIC). Furthermore, bootstrap support for the ML topology was significantly greater when protein rates were used, and some evident errors in the concatenated ML tree topology (i.e., without protein rates) were corrected. [Bayesian credible intervals; DistR method; multigene phylogeny; PHYML; rate heterogeneity.].

Algorithms↗

Finding and fixing systems weaknesses: probabilistic methods and applications of engineering risk analysis.

Methods of engineering risk analysis are based on a functional analysis of systems and on the probabilities (generally Bayesian) of the events and random variables that affect their performances. These methods allow identification of a system's failure modes, computation of its probability of failure or performance deterioration per time unit or operation, and of the contribution of each component to the probabilities and consequences of failures. The model has been extended to include the human decisions and actions that affect components' performances, and the management factors that affect behaviors and can thus be root causes of system failures. By computing the risk with and without proposed measures, one can then set priorities among different risk management options under resource constraints. In this article, I present briefly the engineering risk analysis method, then several illustrations of risk computations that can be used to identify a system's weaknesses and the most cost-effective way to fix them. The first example concerns the heat shield of the space shuttle orbiter and shows the relative risk contribution of the tiles in different areas of the orbiter's surface. The second application is to patient risk in anesthesia and demonstrates how the engineering risk analysis method can be used in the medical domain to rank the benefits of risk mitigation measures, in that case, mostly organizational. The third application is a model of seismic risk analysis and mitigation, with application to the San Francisco Bay area for the assessment of the costs and benefits of different seismic provisions of building codes. In all three cases, some aspects of the results were not intuitively obvious. The probabilistic risk analysis (PRA) method allowed identifying system weaknesses and the most cost-effective way to fix them.

Journal Article↗

Classification of audiograms by sequential testing using a dynamic Bayesian procedure.

A new method for estimating audiograms using behavioral responses is presented. The method is based upon a modification of the Bayesian probability formula in which an outcome is predicted from a static set of events. In the new method, classification of audiograms by sequential testing (CAST), the probabilities of occurrence of audiogram patterns are dynamically updated according to the outcome of each test trial. Computer simulation using an infant response model suggests that the procedure is efficient, sensitive, and specific.

Algorithms↗

Segmenting eukaryotic genomes with the Generalized Gibbs Sampler.

Eukaryotic genomes display segmental patterns of variation in various properties, including GC content and degree of evolutionary conservation. DNA segmentation algorithms are aimed at identifying statistically significant boundaries between such segments. Such algorithms may provide a means of discovering new classes of functional elements in eukaryotic genomes. This paper presents a model and an algorithm for Bayesian DNA segmentation and considers the feasibility of using it to segment whole eukaryotic genomes. The algorithm is tested on a range of simulated and real DNA sequences, and the following conclusions are drawn. Firstly, the algorithm correctly identifies non-segmented sequence, and can thus be used to reject the null hypothesis of uniformity in the property of interest. Secondly, estimates of the number and locations of change-points produced by the algorithm are robust to variations in algorithm parameters and initial starting conditions and correspond to real features in the data. Thirdly, the algorithm is successfully used to segment human chromosome 1 according to GC content, thus demonstrating the feasibility of Bayesian segmentation of eukaryotic genomes. The software described in this paper is available from the author's website (www.uq.edu.au/ approximately uqjkeith/) or upon request to the author.

Algorithms↗

The signed two-space proximity model for learning representations in protein-protein interaction networks.

MOTIVATION: Accurately predicting complex protein-protein interactions (PPIs) is crucial for decoding biological processes, from cellular functioning to disease mechanisms. However, experimental methods for determining PPIs are computationally expensive. Thus, attention has been recently drawn to machine learning approaches. Furthermore, insufficient effort has been made toward analyzing signed PPI networks, which capture both activating (positive) and inhibitory (negative) interactions. To accurately represent biological relationships, we present the Signed Two-Space Proximity Model (S2-SPM) for signed PPI networks, which explicitly incorporates both types of interactions, reflecting the complex regulatory mechanisms within biological systems. This is achieved by leveraging two independent latent spaces to differentiate between positive and negative interactions while representing protein similarity through proximity in these spaces. Our approach also enables the identification of archetypes representing extreme protein profiles. RESULTS: S2-SPM's superior performance in predicting the presence and sign of interactions in SPPI networks is demonstrated in link prediction tasks against relevant baseline methods. Additionally, the biological prevalence of the identified archetypes is confirmed by an enrichment analysis of Gene Ontology (GO) terms, which reveals that distinct biological tasks are associated with archetypal groups formed by both interactions. This study is also validated regarding statistical significance and sensitivity analysis, providing insights into the functional roles of different interaction types. Finally, the robustness and consistency of the extracted archetype structures are confirmed using the Bayesian Normalized Mutual Information (BNMI) metric, proving the model's reliability in capturing meaningful SPPI patterns. AVAILABILITY: S2-SPM is implemented and freely available under the MIT license at https://github.com/Nicknakis/S2SPM.

Protein Interaction Mapping↗

Improving decisionmaking processes with the fuzzy logic approach in the epidemiology of sleep disorders.

Epidemiological studies can provide information not only on specific diagnostic entities but also on their underlying symptomatic constellations. For this purpose, an expert system was developed for the assessment of sleep disorders and endowed with the fuzzy logic capabilities necessary to determine the degree to which a given symptom corresponds to a specific diagnosis. Uncertainty is inherent in fields such as sleep medicine and psychiatry, and becomes evident in clinical practice at the stages of data collection and diagnostic formulation, when the clinician must determine whether a symptom is present and must choose from several diagnostic possibilities. The process involves a considerable degree of subjectivity on the part of the patient in trying to describe his or her symptoms, and of the clinician whose final diagnosis will depend on his or her clinical experience and interpretation of what is normal and what is pathological. Inferential models of the probabilistic or fuzzy logic type take into account such uncertainty. The Sleep-Eval system has been used in epidemiological and clinical studies involving 34,044 interviews collected by close to 300 interviewers. The diagnostic potential of these models is illustrated using data collected in an epidemiological study of the noninstitutionalized general population of Italy and underlines the advantages and limits of the binary, bayesian, and fuzzy logic methods and analyses.

Diagnosis, Computer-Assisted↗

A Bayesian framework for multivariate differential analysis.

Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a hierarchical formulation with moderated-variance estimators for each analyte directly injected into the t-statistic. While standard hypothesis testing strategies are recognised for their low computational cost, allowing for quick extraction of the most differential among thousands of elements, they generally overlook key aspects such as handling missing values, inter-element correlations, and uncertainty quantification. The present paper proposes a fully Bayesian framework for differential analysis, leveraging a conjugate hierarchical formulation for both the mean and the variance. Inference is performed by computing the posterior distribution of compared experimental conditions and sampling from the distribution of differences. This approach provides well-calibrated uncertainty quantification at a similar computational cost as hypothesis testing by leveraging closed-form equations. Furthermore, a natural extension enables multivariate differential analysis that accounts for possible inter-element correlations. We also demonstrate that, in this Bayesian treatment, missing at random data should generally be ignored in univariate settings, and further derive a tailored approximation that handles multiple imputation for the multivariate setting. We argue that probabilistic statements in terms of effect size and associated uncertainty are better suited to practical decision-making. Therefore, we finally propose simple and intuitive inference criteria, such as the overlap coefficient, which express group similarity as a probability rather than traditional, and often misleading, p-values. The performance of this approach is evaluated through an extensive empirical study using both synthetic and controlled real-world proteomics datasets. Overall, we believe that this Bayesian framework for (multivariate) differential analysis provides a valuable and intuitive counterpart to standard methods at a comparable computational cost.

Bayes Theorem↗

From NMR chemical shifts to amino acid types: investigation of the predictive power carried by nuclei.

An approach to automatic prediction of the amino acid type from NMR chemical shift values of its nuclei is presented here, in the frame of a model to calculate the probability of an amino acid type given the set of chemical shifts. The method relies on systematic use of all chemical shift values contained in the BioMagResBank (BMRB). Two programs were designed, one (BMRB stats) for extracting statistical chemical shift parameters from the BMRB and another one (RESCUE2) for computing the probabilities of each amino acid type, given a set of chemical shifts. The Bayesian prediction scheme presented here is compared to other methods already proposed: PROTYP RESCUE and PLATON and is found to be more sensitive and more specific. Using this scheme, we tested various sets of nuclei. The two nuclei carrying the most information are C(beta) and H(beta), in agreement with observations made in Grzesiek and Bax, 1993. Based on four nuclei: H(beta), C(beta), C(alpha) and C', it is possible to increase correct predictions to a rate of more than 75%. Taking into account the correlations between the nuclei chemical shifts has only a slight impact on the percentage of correct predictions: indeed, the largest correlation coefficients display similar features on all amino acids.

Amino Acids↗

Estimating measures of diagnostic accuracy when some covariate information is missing.

Many biomedical data sets are concerned with relating the result of screening procedure(s) for a clinical event to the occurrence of that event. The effect of risk factors on measures of accuracy such as positive predictive value and negative predictive value is of great interest for clinicians. In this paper we propose a generic approach to estimate these measures of accuracy in the setting where an explanatory model has been fitted to the joint screening and event outcome data but information on one or more risk factors in the model is not available. We refer to these as conditional rates, i.e. rates conditioned on only a subset of risk factors. We argue that, based upon the joint distribution of the event outcome, the screening result and the risk factor occurrence, a formal expression for such a rate can be obtained. This expression is a function of model parameters and thus can be estimated once the model has been fitted. Inference within the Bayesian framework is particularly attractive since simulation based model fitting straightforwardly yields samples from the posterior distribution of any conditional rate of interest. We perform a simulation study to compare these estimated conditional rates with frequently used ad hoc estimates. Differences can be substantial. We also illustrate the proposed methodology to compute conditional positive predictive value for a screening mammography data set. The proposed approach is also applicable when there are multiple diagnostic screening test outcomes.

Bayes Theorem↗

Bayesian sample size calculations in phase II clinical trials using a mixture of informative priors.

A number of researchers have discussed phase II clinical trials from a Bayesian perspective. A recent article by Mayo and Gajewski focuses on sample size calculations, which they determine by specifying an informative prior distribution and then calculating a posterior probability that the true response will exceed a prespecified target. In this article, we extend these sample size calculations to include a mixture of informative prior distributions. The mixture comes from several sources of information. For example consider information from two (or more) clinicians. The first clinician is pessimistic about the drug and the second clinician is optimistic. We tabulate the results for sample size design using the fact that the simple mixture of Betas is a conjugate family for the Beta- Binomial model. We discuss the theoretical framework for these types of Bayesian designs and show that the Bayesian designs in this paper approximate this theoretical framework.

Algorithms↗

Theory and application of the maximum likelihood principle to NMR parameter estimation of multidimensional NMR data.

A general theory has been developed for the application of the maximum likelihood (ML) principle to the estimation of NMR parameters (frequency and amplitudes) from multidimensional time-domain NMR data. A computer program (ChiFit) has been written that carries out ML parameter estimation in the D-1 indirectly detected dimensions of a D-dimensional NMR data set. The performance of this algorithm has been tested with experimental three-dimensional (HNCO) and four-dimensional (HN(CO)-CAHA) data from a small protein labeled with 13C and 15N. These data sets, with different levels of digital resolution, were processed using ChiFit for ML analysis and employing conventional Fourier transform methods with prior extrapolation of the time-domain dimensions by linear prediction. Comparison of the results indicates that the ML approach provides superior frequency resolution compared to conventional methods, particularly under conditions of limited digital resolution in the time-domain input data, as is characteristic of D-dimensional NMR data of biomolecules. Close correspondence is demonstrated between the results of analyzing multidimensional time-domain NMR data by Fourier transformation, Bayesian probability theory [Chylla, R.A. and Markley, J.L. (1993) J. Biomol. NMR, 3, 515-533], and the ML principle.

Algorithms↗