Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

Encoding of motion targets by waves in turtle visual cortex.

Visual stimuli evoke wave activity in the visual cortex of freshwater turtles. Earlier work from our laboratory showed that information about the positions of stationary visual stimuli is encoded in the spatiotemporal dynamics of the waves and that the waves can be decoded using Bayesian detection theory. This paper extends these results in three ways. First, it shows that flashes of light separated in space and time and stimuli moving with three speeds can be discriminated statistically using the waves generated in a large-scale model of the cortex. Second, it compares the coding capabilities of spike rate and spike time codes. Spike rate codes were obtained by low-pass filtering the activities of individual neurons in the model with filters of different band widths. For the moving targets used in the study, detectability using spike rate codes is immune to the choice of a specific bandwidth, indicating that a coarse filter is able to adequately discriminate targets. Spike timing codes are binary sequences indicating the precise timing of spike activity of individual neurons across the cortex. Spike time codes generally perform better than do spike rate codes. Third, the encoding process is examined in terms of the underlying cellular mechanisms that result in the initiation, propagation and cessation of the wave. The period of peak detectability corresponds to the period in which waves are propagating across the cortex.

Action Potentials↗

Methods for combining experts' probability assessments.

This article reviews statistical techniques for combining multiple probability distributions. The framework is that of a decision maker who consults several experts regarding some events. The experts express their opinions in the form of probability distributions. The decision maker must aggregate the experts' distributions into a single distribution that can be used for decision making. Two classes of aggregation methods are reviewed. When using a supra Bayesian procedure, the decision maker treats the expert opinions as data that may be combined with its own prior distribution via Bayes' rule. When using a linear opinion pool, the decision maker forms a linear combination of the expert opinions. The major feature that makes the aggregation of expert opinions difficult is the high correlation or dependence that typically occurs among these opinions. A theme of this paper is the need for training procedures that result in experts with relatively independent opinions or for aggregation methods that implicitly or explicitly model the dependence among the experts. Analyses are presented that show that m dependent experts are worth the same as k independent experts where k < or = m. In some cases, an exact value for k can be given; in other cases, lower and upper bounds can be placed on k.

Bayes Theorem↗

Sensitivity of ecological models to their climate drivers: statistical ensembles for forcing.

Global and regional numerical models for terrestrial ecosystem dynamics require fine spatial resolution and temporally complete historical climate fields as input variables. However, because climate observations are unevenly spaced and have incomplete records, such fields need to be estimated. In addition, uncertainty in these fields associated with their estimation are rarely assessed. Ecological models are usually driven with a geostatistical model's mean estimate (kriging) of these fields without accounting for this uncertainty, much less evaluating such errors in terms of their propagation in ecological simulations. We introduce a Bayesian statistical framework to model climate observations to create spatially uniform and temporally complete fields, taking into account correlation in time and space, spatial heterogeneity, lack of normality, and uncertainty about all these factors. A key benefit of the Bayesian model is that it generates uncertainty measures for the generated fields. To demonstrate this method, we reconstruct historical monthly precipitation fields (a driver for ecological models) on a fine resolution grid for a climatically heterogeneous region in the western United States. The main goal of this work is to evaluate the sensitivity of ecological models to the uncertainty associated with prediction of their climate drivers. To assess their numerical sensitivity to predicted input variables, we generate a set of ecological model simulations run using an ensemble of different versions of the reconstructed fields. We construct such an ensemble by sampling from the posterior predictive distribution of the climate field. We demonstrate that the estimated prediction error of the climate field can be very high. We evaluate the importance of such errors in ecological model experiments using an ensemble of historical precipitation time series in simulations of grassland biogeochemical dynamics with an ecological numerical model, Century. We show how uncertainty in predicted precipitation fields is propagated into ecological model results and that this propagation had different modes. Depending on output variable, the response of model dynamics to uncertainty in inputs ranged from uncertainty in outputs that matched that of inputs to those that were muted or that were biased, as well as uncertainty that was persistent in time after input errors dropped.

Animals↗

Image construction methods for phased array magnetic resonance imaging.

PURPOSE: To study image construction in phased array magnetic resonance imaging (MRI) systems from a statistical signal processing point of view. MATERIALS AND METHODS: Three new approaches for image combination with multiple coils are proposed: 1) one based on the singular value decomposition of the measurement matrix, which is asymptotically optimal in the signal-to-noise ratio sense; 2) one based on a maximum-likelihood formulation, incorporating a priori information on the coil sensitivities in a Bayesian manner; and 3) one based on a least-squares formulation, which incorporates a smoothness constraint on the coil sensitivities. RESULTS: Numerical examples using synthetic and real data are presented to illustrate the performance of these new approaches. Results on the synthetic data show improvement in signal-to-error ratio, while results on the real data (a 4.7 T four-coil image of a cat spinal cord) show that the proposed methods can improve the SNR in the final image by up to 3 dB in the regions of interest compared to conventional sum-of-squares processing. CONCLUSION: It is demonstrated that phased array MRI reconstruction performance can be improved by the use of more elaborate statistical signal processing algorithms.

Algorithms↗

Bayesian pharmacokinetics of lithium after an acute self-intoxication and subsequent haemodialysis: a case report.

We report a case of a 39-year-old male with bipolar affective disorder who was admitted to hospital with an intentional acute lithium intoxication resulting in renal insufficiency. The patient had previously been treated with lithium, risperidone, fluoxetine and lorazepam, and successfully titrated to lithium levels of 0.7 mmol/l. After overdosing, the lithium level was 5.89 mmol/l and haemodialysis was initiated. A full pharmacokinetic time profile of lithium was obtained. After successful haemodialysis treatment, lithium levels recovered below toxic levels of 1.5 mmol/l in 53 hr. Without intervention non-toxic levels were not expected to have been reached within 6 days, based on computer simulation of predialysis levels. The patient was discharged 6 days after admission without residual symptoms. It was concluded that the lithium intoxication resulted from a combination of lithium overdose and subsequent renal insufficiency due to the overdose. A possible fluoxetine-risperidone interaction was not considered clinically apparent.

Acute Disease↗

Precaution, uncertainty and causation in environmental decisions.

What measures of uncertainty and what causal analysis can improve the management of potentially severe, irreversible or dreaded environmental outcomes? Environmental choices show that policies intended to be precautionary (such as adding MTBE to petrol) can cause unanticipated harm (by mobilizing benzene, a known leukemogen, in the ground water). Many environmental law principles set the boundaries of what should be done but do not provide an operational construct to answer this question. Those principles, ranging from the precautionary principle to protecting human health from a significant risk of material health impairment, do not explain how to make environmental management choices when incomplete, inconsistent and complex scientific evidence characterizes potentially adverse environmental outcomes. Rather, they pass the task to lower jurisdictions such as agencies or authorities. To achieve the goals of the principle, those who draft it must deal with scientific casual conjectures, partial knowledge and variable data. In this paper we specifically deal with the qualitative and quantitative aspects of the European Union's (EU) explanation of consistency and on the examination of scientific developments relevant to variability and uncertain data and causation. Managing hazards under the precautionary principle requires inductive, empirical methods of assessment. However, acting on a scientific conjecture can also be socially unfair, costly, and detrimental when applied to complex environmental choices. We describe a constructive framework rationally to meet the command of the precautionary principle using alternative measures of uncertainty and recent statistical methods of causal analysis. These measures and methods can bridge the gap between conjectured future irreversible or severe harm and scant scientific evidence, thus leading to more confident and resilient social choices. We review two sets of measures and computational systems to deal with uncertainty and link them to causation through inductive empirical methods such as Bayesian Networks. We conclude that primary legislation concerned with large uncertainties and potential severe or dreaded environmental outcomes can produce accurate and efficient choices. To do so, primary legislation should specifically indicate what measures can represent uncertainty and how to deal with uncertain causation thus providing guidance to an agency's rulemaking or to an authority's writing secondary legislation. A corollary conclusion with legal, scientific and probabilistic implications concerns how to update past information when the state of information increases because a failure to update can result in regretting past choices. Elected legislators have the democratic mandate to formulate precautionary principles and are accountable. To preserve that mandate, imbedding formal methods to represent uncertainty in the statutory language of the precautionary principle enhances subsequent judicial review of legislative actions. The framework that we propose also reduces the Balkanized views and interpretations of probabilities, possibilities, likelihood and uncertainty that exists in environmental decision-making.

Bayes Theorem↗

BYY harmony learning, structural RPCL, and topological self-organizing on mixture models.

The Bayesian Ying-Yang (BYY) harmony learning acts as a general statistical learning framework, featured by not only new regularization techniques for parameter learning but also a new mechanism that implements model selection either automatically during parameter learning or via a new class of model selection criteria used after parameter learning. In this paper, further advances on BYY harmony learning by considering modular inner representations are presented in three parts. One consists of results on unsupervisedmixture models, ranging from Gaussian mixture based Mean Square Error (MSE) clustering, elliptic clustering, subspace clustering to NonGaussian mixture based clustering not only with each cluster represented via either Bernoulli-Gaussian mixtures or independent real factor models, but also with independent component analysis implicitly made on each cluster. The second consists of results on supervised mixture-of-experts (ME) models, including Gaussian ME, Radial Basis Function nets, and Kernel regressions. The third consists of two strategies for extending the above structural mixtures into self-organized topological maps. All these advances are introduced with details on three issues, namely, (a) adaptive learning algorithms, especially elliptic, subspace, and structural rival penalized competitive learning algorithms, with model selection made automatically during learning; (b) model selection criteria for being used after parameter learning, and (c) how these learning algorithms and criteria are obtained from typical special cases of BYY harmony learning.

Bayes Theorem↗

Bayesian infinite mixture model based clustering of gene expression profiles.

MOTIVATION: The biologic significance of results obtained through cluster analyses of gene expression data generated in microarray experiments have been demonstrated in many studies. In this article we focus on the development of a clustering procedure based on the concept of Bayesian model-averaging and a precise statistical model of expression data. RESULTS: We developed a clustering procedure based on the Bayesian infinite mixture model and applied it to clustering gene expression profiles. Clusters of genes with similar expression patterns are identified from the posterior distribution of clusterings defined implicitly by the stochastic data-generation model. The posterior distribution of clusterings is estimated by a Gibbs sampler. We summarized the posterior distribution of clusterings by calculating posterior pairwise probabilities of co-expression and used the complete linkage principle to create clusters. This approach has several advantages over usual clustering procedures. The analysis allows for incorporation of a reasonable probabilistic model for generating data. The method does not require specifying the number of clusters and resulting optimal clustering is obtained by averaging over models with all possible numbers of clusters. Expression profiles that are not similar to any other profile are automatically detected, the method incorporates experimental replicates, and it can be extended to accommodate missing data. This approach represents a qualitative shift in the model-based cluster analysis of expression data because it allows for incorporation of uncertainties involved in the model selection in the final assessment of confidence in similarities of expression profiles. We also demonstrated the importance of incorporating the information on experimental variability into the clustering model. AVAILABILITY: The MS Windows(TM) based program implementing the Gibbs sampler and supplemental material is available at http://homepages.uc.edu/~medvedm/BioinformaticsSupplement.htm CONTACT: medvedm@email.uc.edu

Bayes Theorem↗

A Markov model for blind image separation by a mean-field EM algorithm.

This paper deals with blind separation of images from noisy linear mixtures with unknown coefficients, formulated as a Bayesian estimation problem. This is a flexible framework, where any kind of prior knowledge about the source images and the mixing matrix can be accounted for. In particular, we describe local correlation within the individual images through the use of Markov random field (MRF) image models. These are naturally suited to express the joint pdf of the sources in a factorized form, so that the statistical independence requirements of most independent component analysis approaches to blind source separation are retained. Our model also includes edge variables to preserve intensity discontinuities. MRF models have been proved to be very efficient in many visual reconstruction problems, such as blind image restoration, and allow separation and edge detection to be performed simultaneously. We propose an expectation-maximization algorithm with the mean field approximation to derive a procedure for estimating the mixing matrix, the sources, and their edge maps. We tested this procedure on both synthetic and real images, in the fully blind case (i.e., no prior information on mixing is exploited) and found that a source model accounting for local autocorrelation is able to increase robustness against noise, even space variant. Furthermore, when the model closely fits the source characteristics, independence is no longer a strict requirement, and cross-correlated sources can be separated, as well.

Algorithms↗

A Bayesian model for triage decision support.

OBJECTIVE: To compare triage decisions of an automated emergency department triage system with decisions made by an emergency specialist. METHODS: In a retrospective setting, data extracted from charts of 90 patients with chief complaint of non-traumatic abdominal pain were used as input for triage system and emergency medicine specialist. The final disposition and diagnoses of the physicians who visited the patient in Emergency Department (ED) as reflected in the medical records were considered as control. Results were compared by chi(2)-test and a binary logistic regression model. RESULTS: Compared to emergency medicine specialist, triage system had higher sensitivity (90% versus 64%) and lower specificity (25% versus 48%) for patients who required hospitalization. The triage system successfully predicted the Admit decisions made in the ED whereas the emergency medicine specialist decisions could not predict the ED disposition. Both triage system and emergency medicine specialist properly disposed 56% of cases, however, the emergency medicine specialist in this study under-disposed more patients than the triage system considering Admit disposition (p=0.004) while he appropriately discharged more patients compared to the triage system (p=0.017). CONCLUSION: The triage system studied here shows promise as a triage decision support tool to be used for telephone triage and triage in the emergency departments. This technology may also be useful to the patients as a self-triage tool. However, the efficiency of this particular application of this technology is unclear.

Abdominal Pain↗

Estimating haplotype frequencies and standard errors for multiple single nucleotide polymorphisms.

Estimating haplotype frequencies becomes increasingly important in the mapping of complex disease genes, as millions of single nucleotide polymorphisms (SNPs) are being identified and genotyped. When genotypes at multiple SNP loci are gathered from unrelated individuals, haplotype frequencies can be accurately estimated using expectation-maximization (EM) algorithms (Excoffier and Slatkin, 1995; Hawley and Kidd, 1995; Long et al., 1995), with standard errors estimated using bootstraps. However, because the number of possible haplotypes increases exponentially with the number of SNPs, handling data with a large number of SNPs poses a computational challenge for the EM methods and for other haplotype inference methods. To solve this problem, Niu and colleagues, in their Bayesian haplotype inference paper (Niu et al., 2002), introduced a computational algorithm called progressive ligation (PL). But their Bayesian method has a limitation on the number of subjects (no more than 100 subjects in the current implementation of the method). In this paper, we propose a new method in which we use the same likelihood formulation as in Excoffier and Slatkin's EM algorithm and apply the estimating equation idea and the PL computational algorithm with some modifications. Our proposed method can handle data sets with large number of SNPs as well as large numbers of subjects. Simultaneously, our method estimates standard errors efficiently, using the sandwich-estimate from the estimating equation, rather than the bootstrap method. Additionally, our method admits missing data and produces valid estimates of parameters and their standard errors under the assumption that the missing genotypes are missing at random in the sense defined by Rubin (1976).

Algorithms↗

Using sensitivity analysis for efficient quantification of a belief network.

Sensitivity analysis is a method to investigate the effects of varying a model's parameters on its predictions. It was recently suggested as a suitable means to facilitate quantifying the joint probability distribution of a Bayesian belief network. This article presents practical experience with performing sensitivity analyses on a belief network in the field of medical prognosis and treatment planning. Three network quantifications with different levels of informedness were constructed. Two poorly-informed quantifications were improved by replacing the most influential parameters with the corresponding parameter estimates from the well-informed network quantification; these influential parameters were found by performing one-way sensitivity analyses. Subsequently, the results of the replacements were investigated by comparing network predictions. It was found that it may be sufficient to gather a limited number of highly-informed network parameters to obtain a satisfying network quantification. It is therefore concluded that sensitivity analysis can be used to improve the efficiency of quantifying a belief network.

Algorithms↗

Recursive noisy OR--a rule for estimating complex probabilistic interactions.

This paper focuses on approaches that address the intractability of knowledge acquisition of conditional probability tables in causal or Bayesian belief networks. We state a rule that we term the "recursive noisy OR" (RNOR) which allows combinations of dependent causes to be entered and later used for estimating the probability of an effect. In the development of this paper, we investigate the axiomatic correctness and semantic meaning of this rule and show that the recursive noisy OR is a generalization of the well-known noisy OR. We introduce the concept of positive causality and demonstrate its utility in axiomatic correctness of the RNOR. We also introduce concepts describing the ways in which dependent causes can work together as being either "synergistic" or "interfering." We provide a formalization to quantify these concepts and show that they are preserved by the RNOR. Finally, we present a method for the determination of Conditional Probability Tables from this causal theory.

Algorithms↗

Robust and accurate Bayesian inference of genome-wide genealogies for hundreds of genomes.

The Ancestral Recombination Graph (ARG), which describes the genealogical history of a sample of genomes, is a vital tool in population genomics and biomedical research. Recent advancements have substantially increased ARG reconstruction scalability, but they rely on approximations that can reduce accuracy, especially under model misspecification. Moreover, they reconstruct only a single ARG topology and cannot quantify the considerable uncertainty associated with ARG inferences. Here, to address these challenges, we introduce SINGER (sampling and inferring of genealogies with recombination), a method that accelerates ARG sampling from the posterior distribution by two orders of magnitude, enabling accurate inference and uncertainty quantification for hundreds of whole-genome sequences. Through extensive simulations, we demonstrate SINGER's enhanced accuracy and robustness to model misspecification compared to existing methods. We demonstrate the utility of SINGER by applying it to individuals of British and African descent within the 1000 Genomes Project, identifying signals of population differentiation, archaic introgression and strong support for ancient polymorphism in the human leukocyte antigen region shared across primates.

Humans↗

Detecting Interspecific Positive Selection Using Convolutional Neural Networks.

Traditional statistical methods using maximum likelihood and Bayesian inference can detect positive selection from an interspecific phylogeny and a codon sequence alignment based on model assumptions, but they are prone to false positives due to alignment errors and can lack power. These problems are particularly pronounced when faced with high levels of indels and divergence. To address these issues, we trained and tested convolutional neural network models on simulated data and achieved higher accuracy in detecting selection across a specific range of phylogenetic scenarios and evolutionary modes. This advantage is particularly evident when performing inference on noisy data prone to misalignments. Our method shows some ability to account for these errors, where most statistical frameworks fail to do so in a tractable manner. We explore the generalizability of our convolutional neural network models to unseen evolutionary scenarios and identify future avenues to achieve broader utility. Once trained, our convolutional neural network model is faster at test time, making it a scalable alternative to traditional statistical methods for large-scale, multigene analyses. In addition to binary classification (inference of the presence or absence of positive selection during the evolution of the sequences), we use saliency maps to understand what the model learns and observe how this could be leveraged for sitewise inference of positive selection.

Neural Networks, Computer↗

Protein molecular function prediction by Bayesian phylogenomics.

We present a statistical graphical model to infer specific molecular function for unannotated protein sequences using homology. Based on phylogenomic principles, SIFTER (Statistical Inference of Function Through Evolutionary Relationships) accurately predicts molecular function for members of a protein family given a reconciled phylogeny and available function annotations, even when the data are sparse or noisy. Our method produced specific and consistent molecular function predictions across 100 Pfam families in comparison to the Gene Ontology annotation database, BLAST, GOtcha, and Orthostrapper. We performed a more detailed exploration of functional predictions on the adenosine-5'-monophosphate/adenosine deaminase family and the lactate/malate dehydrogenase family, in the former case comparing the predictions against a gold standard set of published functional characterizations. Given function annotations for 3% of the proteins in the deaminase family, SIFTER achieves 96% accuracy in predicting molecular function for experimentally characterized proteins as reported in the literature. The accuracy of SIFTER on this dataset is a significant improvement over other currently available methods such as BLAST (75%), GeneQuiz (64%), GOtcha (89%), and Orthostrapper (11%). We also experimentally characterized the adenosine deaminase from Plasmodium falciparum, confirming SIFTER's prediction. The results illustrate the predictive power of exploiting a statistical model of function evolution in phylogenomic problems. A software implementation of SIFTER is available from the authors.

Adenosine Deaminase↗

Designing libraries with CNS activity.

Library design is an important and difficult task. In this paper we describe one possible solution to designing a CNS-active library. CNS-actives and -inactives were selected from the CMC and the MDDR databases based on whether they were described as having some kind of CNS activity in the databases. This classification scheme results in over 15 000 actives and over 50 000 inactives. Each molecule is described by 7 1D descriptors (molecular weight, number of donors, number of acceptors, etc.) and 166 2D descriptors (presence/absence of functional groups such as NH(2)). A neural network trained using Bayesian methods can correctly predict about 75% of the actives and 65% of the inactives using the 7 1D descriptors. The performance improves to a prediction accuracy on the active set of 83% and 79% on the inactives on adding the 2D descriptors. On a database with 275 compounds where the CNS activity is known (from the literature) for each compound, we achieve 92% and 71% accuracy on the actives and inactives, respectively. The models we construct can therefore be used as a "filter" to examine any set of proposed molecules in a chemical library. As an example of the utility of our method, we describe the generation of a small library of potentially CNS-active molecules that would be amenable to combinatorial chemistry. This was done by building and analyzing a large database of a million compounds constructed from frameworks and side chains frequently found in drug molecules.

Animals↗

Collateral missing value imputation: a new robust missing value estimation algorithm for microarray data.

MOTIVATION: Microarray data are used in a range of application areas in biology, although often it contains considerable numbers of missing values. These missing values can significantly affect subsequent statistical analysis and machine learning algorithms so there is a strong motivation to estimate these values as accurately as possible before using these algorithms. While many imputation algorithms have been proposed, more robust techniques need to be developed so that further analysis of biological data can be accurately undertaken. In this paper, an innovative missing value imputation algorithm called collateral missing value estimation (CMVE) is presented which uses multiple covariance-based imputation matrices for the final prediction of missing values. The matrices are computed and optimized using least square regression and linear programming methods. RESULTS: The new CMVE algorithm has been compared with existing estimation techniques including Bayesian principal component analysis imputation (BPCA), least square impute (LSImpute) and K-nearest neighbour (KNN). All these methods were rigorously tested to estimate missing values in three separate non-time series (ovarian cancer based) and one time series (yeast sporulation) dataset. Each method was quantitatively analyzed using the normalized root mean square (NRMS) error measure, covering a wide range of randomly introduced missing value probabilities from 0.01 to 0.2. Experiments were also undertaken on the yeast dataset, which comprised 1.7% actual missing values, to test the hypothesis that CMVE performed better not only for randomly occurring but also for a real distribution of missing values. The results confirmed that CMVE consistently demonstrated superior and robust estimation capability of missing values compared with other methods for both series types of data, for the same order of computational complexity. A concise theoretical framework has also been formulated to validate the improved performance of the CMVE algorithm. AVAILABILITY: The CMVE software is available upon request from the authors.

Algorithms↗