Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

Joint Bayesian estimation of alignment and phylogeny.

We describe a novel model and algorithm for simultaneously estimating multiple molecular sequence alignments and the phylogenetic trees that relate the sequences. Unlike current techniques that base phylogeny estimates on a single estimate of the alignment, we take alignment uncertainty into account by considering all possible alignments. Furthermore, because the alignment and phylogeny are constructed simultaneously, a guide tree is not needed. This sidesteps the problem in which alignments created by progressive alignment are biased toward the guide tree used to generate them. Joint estimation also allows us to model rate variation between sites when estimating the alignment and to use the evidence in shared insertion/deletions (indels) to group sister taxa in the phylogeny. Our indel model makes use of affine gap penalties and considers indels of multiple letters. We make the simplifying assumption that the indel process is identical on all branches. As a result, the probability of a gap is independent of branch length. We use a Markov chain Monte Carlo (MCMC) method to sample from the posterior of the joint model, estimating the most probable alignment and tree and their support simultaneously. We describe a new MCMC transition kernel that improves our algorithm's mixing efficiency, allowing the MCMC chains to converge even when started from arbitrary alignments. Our software implementation can estimate alignment uncertainty and we describe a method for summarizing this uncertainty in a single plot.

Algorithms↗

[Bayesian representation of prior information and MCMC method in microwave imaging].

Microwave imaging for dielectric objects was considered in this paper. Applying Bayesian approach to represent prior information about permittivity distribution of observed object by prior probability density and combine measurements information of scattering field, we obtained posterior probability density that included synthetic information about the observed object. And then, Gibbs sampler, one of Markov Chain Monte Carlo method, was used to sample the posterior probability density. The sample mean was regarded as an evaluation of the permittivity distribution. The results of simulation imaging with "blocky" objects showed that this set of methods made good use of information and had the advantages of feasibility and very strong anti-noise ability. In addition,it is capable of describing (definite or indefinite) prior information in a convenient and controllable way, as well as capable of giving the "complete" solution, i.e., the occurrence probability of every permittivity distribution.

Bayes Theorem↗

Integrated surface model optimization for freehand three-dimensional echocardiography.

The major obstacle of three-dimensional (3-D) echocardiography is that the ultrasound image quality is too low to reliably detect features locally. Almost all available surface-finding algorithms depend on decent quality boundaries to get satisfactory surface models. We formulate the surface model optimization problem in a Bayesian framework, such that the inference made about a surface model is based on the integration of both the low-level image evidence and the high-level prior shape knowledge through a pixel class prediction mechanism. We model the probability of pixel classes instead of making explicit decisions about them. Therefore, we avoid the unreliable edge detection or image segmentation problem and the pixel correspondence problem. An optimal surface model best explains the observed images such that the posterior probability of the surface model for the observed images is maximized. The pixel feature vector as the image evidence includes several parameters such as the smoothed grayscale value and the minimal second directional derivative. Statistically, we describe the feature vector by the pixel appearance probability model obtained by a nonparametric optimal quantization technique. Qualitatively, we display the imaging plane intersections of the optimized surface models together with those of the ground-truth surfaces reconstructed from manual delineations. Quantitatively, we measure the projection distance error between the optimized and the ground-truth surfaces. In our experiment, we use 20 studies to obtain the probability models offline. The prior shape knowledge is represented by a catalog of 86 left ventricle surface models. In another set of 25 test studies, the average epicardial and endocardial surface projection distance errors are 3.2 +/- 0.85 mm and 2.6 +/- 0.78 mm, respectively.

Algorithms↗

A comparative investigation on subspace dimension determination.

It is well-known that constrained Hebbian self-organization on multiple linear neural units leads to the same k-dimensional subspace spanned by the first k principal components. Not only the batch PCA algorithm has been widely applied in various fields since 1930s, but also a variety of adaptive algorithms have been proposed in the past two decades. However, most studies assume a known dimension k or determine it heuristically, though there exist a number of model selection criteria in the literature of statistics. Recently, criteria have also been obtained under the framework of Bayesian Ying-Yang (BYY) harmony learning. This paper further investigates the BYY criteria in comparison with existing typical criteria, including Akaike's information criterion (AIC), the consistent Akaike's information criterion (CAIC), the Bayesian inference criterion (BIC), and the cross-validation (CV) criterion. This comparative study is made via experiments not only on simulated data sets of different sample sizes, noise variances, data space dimensions, and subspace dimensions, but also on two real data sets from air pollution problem and sport track records, respectively. Experiments have shown that BIC outperforms AIC, CAIC, and CV while the BYY criteria are either comparable with or better than BIC. Therefore, BYY harmony learning is a more preferred tool for subspace dimension determination by further considering that the appropriate subspace dimension k can be automatically determined during implementing BYY harmony learning for the principal subspace while the selection of subspace dimension k by BIC, AIC, CAIC, and CV has to be made at the second stage based on a set of candidate subspaces with different dimensions which have to be obtained at the first stage of learning.

Air Pollution↗

Bayesian segregation analysis of milk flow in Swiss dairy cattle using Gibbs sampling.

Segregation analyses with Gibbs sampling were applied to investigate the mode of inheritance and to estimate the genetic parameters of milk flow of Swiss dairy cattle. The data consisted of 204,397, 655,989 and 40,242 lactation records of milk flow in Brown Swiss, Simmental and Holstein cattle, respectively (4 to 22 years). Separate genetic analyses of first and multiple lactations were carried out for each breed. The results show that genetic parameters especially polygenic variance and heritability of milk flow in the first lactation were very similar under both mixed inheritance (polygenes + major gene) and polygenic models. Segregation analyses yielded very low major gene variances which favour the polygenic determinism of milk flow. Heritabilities and repeatabilities of milk flow in both Brown Swiss and Simmental were high (0.44 to 0.48 and 0.54 to 0.59, respectively). The heritability of milk flow based on scores of milking ability in Holstein was intermediate (0.25). Variance components and heritabilities in the first lactation were slightly larger than those estimates for multiple lactations. The results suggest that milk flow (the quantity of milk per minute of milking) is a relevant measurement to characterise the cows milking ability which is a good candidate trait to be evaluated for a possible inclusion in the selection objectives in dairy cattle.

Animals↗

Geostatistical analysis of disease data: accounting for spatial support and population density in the isopleth mapping of cancer mortality risk using area-to-point Poisson kriging.

BACKGROUND: Geostatistical techniques that account for spatially varying population sizes and spatial patterns in the filtering of choropleth maps of cancer mortality were recently developed. Their implementation was facilitated by the initial assumption that all geographical units are the same size and shape, which allowed the use of geographic centroids in semivariogram estimation and kriging. Another implicit assumption was that the population at risk is uniformly distributed within each unit. This paper presents a generalization of Poisson kriging whereby the size and shape of administrative units, as well as the population density, is incorporated into the filtering of noisy mortality rates and the creation of isopleth risk maps. An innovative procedure to infer the point-support semivariogram of the risk from aggregated rates (i.e. areal data) is also proposed. RESULTS: The novel methodology is applied to age-adjusted lung and cervix cancer mortality rates recorded for white females in two contrasted county geographies: 1) state of Indiana that consists of 92 counties of fairly similar size and shape, and 2) four states in the Western US (Arizona, California, Nevada and Utah) forming a set of 118 counties that are vastly different geographical units. Area-to-point (ATP) Poisson kriging produces risk surfaces that are less smooth than the maps created by a naïve point kriging of empirical Bayesian smoothed rates. The coherence constraint of ATP kriging also ensures that the population-weighted average of risk estimates within each geographical unit equals the areal data for this unit. Simulation studies showed that the new approach yields more accurate predictions and confidence intervals than point kriging of areal data where all counties are simply collapsed into their respective polygon centroids. Its benefit over point kriging increases as the county geography becomes more heterogeneous. CONCLUSION: A major limitation of choropleth maps is the common biased visual perception that larger rural and sparsely populated areas are of greater importance. The approach presented in this paper allows the continuous mapping of mortality risk, while accounting locally for population density and areal data through the coherence constraint. This form of Poisson kriging will facilitate the analysis of relationships between health data and putative covariates that are typically measured over different spatial supports.

Cluster Analysis↗

Modelling of mortality data from a multi-centre study in Japan by means of Poisson regression with error in variables.

BACKGROUND: Death rates of particular categories in epidemiological studies are often based on a small number of occurrences which can be well described by a Poisson distribution. METHOD: We applied this model for the analysis of a multi-centre study in five Japanese counties where the death rates of stomach cancer (ICD-9 code 151) in four age groups are known. In our example some covariates of the cases (e.g. plasma lycopene levels) are unknown values and are estimated from a randomly chosen collective. Therefore these values are subject to a sampling error. The inclusion of errors in variables (e-i-v) into the statistical model can adequately describe such a situation. The model is estimated in a Bayesian framework by means of resampling techniques. RESULTS: Based on the posterior distribution of the parameters the relative risk of stomach cancer is 0.46 (95% confidence interval: 0.23-0.79) comparing the maximum of the population medians of lycopene with the minimum. The estimated overdispersion is close to zero indicating only minor interference with other possible explanatory variables. In addition, we show that inclusion of e-i-v can give more accurate estimates of the parameters even from small sample sizes. CONCLUSIONS: Appropriate statistical methods allow the accurate estimation of relative risks from small sample sizes and from low number of cases. Lycopene plasma levels are good predictors for stomach cancer.

Adult↗

Bayesian estimation of the timing and severity of a population bottleneck from ancient DNA.

In this first application of the approximate Bayesian computation approach using the serial coalescent, we demonstrated the estimation of historical demographic parameters from ancient DNA. We estimated the timing and severity of a population bottleneck in an endemic subterranean rodent, Ctenomys sociabilis, over the last 10,000 y from two cave sites in northern Patagonia, Argentina. Understanding population bottlenecks is important in both conservation and evolutionary biology. Conservation implications include the maintenance of genetic variation, inbreeding, fixation of mildly deleterious alleles, and loss of adaptive potential. Evolutionary processes are impacted because of the influence of small populations in founder effects and speciation. We found a decrease from a female effective population size of 95,231 to less than 300 females at 2,890 y before present: a 99.7% decline. Our study demonstrates the persistence of a species depauperate in genetic diversity for at least 2,000 y and has implications for modes of speciation in the incredibly diverse rodent genus Ctenomys. Our approach shows promise for determining demographic parameters for other species with ancient and historic samples and demonstrates the power of such an approach using ancient DNA.

Animals↗

Estimating Re and overdispersion in secondary cases from the size of identical sequence clusters of SARS-CoV-2.

The wealth of genomic data that was generated during the COVID-19 pandemic provides an exceptional opportunity to obtain information on the transmission of SARS-CoV-2. Specifically, there is great interest to better understand how the effective reproduction number [Formula: see text] and the overdispersion of secondary cases, which can be quantified by the negative binomial dispersion parameter k, changed over time and across regions and viral variants. The aim of our study was to develop a Bayesian framework to infer [Formula: see text] and k from viral sequence data. First, we developed a mathematical model for the distribution of the size of identical sequence clusters, in which we integrated viral transmission, the mutation rate of the virus, and incomplete case-detection. Second, we implemented this model within a Bayesian inference framework, allowing the estimation of [Formula: see text] and k from genomic data only. We validated this model in a simulation study. Third, we identified clusters of identical sequences in all SARS-CoV-2 sequences in 2021 from Switzerland, Denmark, and Germany that were available on GISAID. We obtained monthly estimates of the posterior distribution of [Formula: see text] and k, with the resulting [Formula: see text] estimates slightly lower than estimates obtained by other methods, and k comparable with previous results. We found comparatively higher estimates of k in Denmark which suggests less opportunities for superspreading and more controlled transmission compared to the other countries in 2021. Our model included an estimation of the case detection and sampling probability, but the estimates obtained had large uncertainty, reflecting the difficulty of estimating these parameters simultaneously. Our study presents a novel method to infer information on the transmission of infectious diseases and its heterogeneity using genomic data. With increasing availability of sequences of pathogens in the future, we expect that our method has the potential to provide new insights into the transmission and the overdispersion in secondary cases of other pathogens.

COVID-19↗

Dynamic causal modelling.

In this paper we present an approach to the identification of nonlinear input-state-output systems. By using a bilinear approximation to the dynamics of interactions among states, the parameters of the implicit causal model reduce to three sets. These comprise (1) parameters that mediate the influence of extrinsic inputs on the states, (2) parameters that mediate intrinsic coupling among the states, and (3) [bilinear] parameters that allow the inputs to modulate that coupling. Identification proceeds in a Bayesian framework given known, deterministic inputs and the observed responses of the system. We developed this approach for the analysis of effective connectivity using experimentally designed inputs and fMRI responses. In this context, the coupling parameters correspond to effective connectivity and the bilinear parameters reflect the changes in connectivity induced by inputs. The ensuing framework allows one to characterise fMRI experiments, conceptually, as an experimental manipulation of integration among brain regions (by contextual or trial-free inputs, like time or attentional set) that is revealed using evoked responses (to perturbations or trial-bound inputs, like stimuli). As with previous analyses of effective connectivity, the focus is on experimentally induced changes in coupling (cf., psychophysiologic interactions). However, unlike previous approaches in neuroimaging, the causal model ascribes responses to designed deterministic inputs, as opposed to treating inputs as unknown and stochastic.

Arousal↗

Estimating transitions between symptom severity states over time in schizophrenia: a Bayesian meta-analytic approach.

We obtain the posterior predictive distribution of transition probabilities between symptom severity states over time for patients with schizophrenia by (i) employing a Bayesian meta-analysis of published clinical trials and observational studies to estimate the posterior distribution of parameters that guide changes in Positive and Negative Syndrome Scale (PANSS) scores over time and under the influence of various drugs and (ii) by propagating the variability from the posterior distributions of the parameters through a micro-simulation model that is formulated based on schizophrenia progression. Results show detailed differences among haloperidol, risperidone and olanzapine in controlling various levels of severities of positive, negative and joint symptoms over time. For example, risperidone seems best in controlling severe positive symptoms while olanzapine is the worst in that during the first quarter of drug treatment; however, olanzapine seems to be best in controlling severe negative symptoms across all four quarters of treatment while haloperidol is the worst in this regard. These details may further serve to better estimate quality of life of patients and aid in resource utilization decisions in treating schizophrenic patients. In addition, consistent estimation of uncertainty in the time-profile parameters also has important implications for the practice of cost-effectiveness analysis and for future resource allocation policies in schizophrenia treatment.

Antipsychotic Agents↗

A Bayesian model for assessing the frequency of multiple mating in nature.

Many breeding systems have multiple mating, in which males or females mate with multiple partners. With the advent of molecular markers, it is now possible to detect multiple mating in nature. However, no model yet exists to effectively assess the frequency of multiple mating (f(mm))--the proportion of broods with at least two males (or females) genetically contributing--from limited genetic data. We present a single-sex model based on Bayes' rule that incorporates the numbers of loci, alleles, offspring, and genetic parents. Two genetic criteria for calculating f(mm) are considered: the proportion of broods with three or more paternal (or maternal) alleles at any one locus and the total number of haplotypes observed in each brood. The former criterion provides the most precise estimates of f(mm). The model enables the calculation of confidence intervals and allows mutations (or typing errors) to be incorporated into the calculation. Failure to account for mutations can result in overestimates of f(mm). The model can also utilize other biological data, such as behavioral observations during mating, thereby increasing the accuracy of the calculation as compared to previous models. For example, when two sires contribute equally to multiply mated broods, only three loci with five equally common alleles are required to provide estimates of f(mm) with high precision. We demonstrate the model with an example addressing the frequency of multiple paternity in small versus large clutches of the endangered Kemp's Ridley sea turtle (Lepidochelys kempi) and show that females that lay large clutches are more likely to have multiply mated.

Animals↗

The continual reassessment method and its applications: a Bayesian methodology for phase I cancer clinical trials.

We discuss the continual reassessment method (CRM) and its extension with practical applications in phase I and I/II cancer clinical trials. The CRM has been proposed as an alternative design of a traditional cohort design and its essential features are the sequential (continual) selection of a dose level for the next patients based on the dose-toxicity relationship and the updating of the relationship based on patients' response data using Bayesian calculation. The original CRM has been criticized because it often tends to allocate too toxic doses to many patients and our proposal for overcoming this practical problem is to monitor a posterior density function of the occurrence of the dose limiting toxicity (DLT) at each dose level. A simulation study shows that strategies based on our proposal allocate a smaller number of patients to doses higher than the maximum tolerated dose (MTD) compared with the original method while the mean squared error of the probability of the DLT occurrence at the MTD is not inflated. We present a couple of extensions of the CRM with real prospective applications: (i) monitoring efficacy and toxicity simultaneously in a combination phase I/II trial; (ii) combining the idea of pharmacokinetically guided dose escalation (PKGDE) and utilization of animal toxicity data in determining the prior distribution. A stopping rule based on the idea of separation among the DLT density functions is discussed in the first example and a strategy for determining the model parameter of the dose-toxicity relationship is suggested in the second example.

Animals↗

Coronary artery bypass risk prediction using neural networks.

BACKGROUND: Neural networks are nonparametric, robust, pattern recognition techniques that can be used to model complex relationships. METHODS: The applicability of multilayer perceptron neural networks (MLP) to coronary artery bypass grafting risk prediction was assessed using The Society of Thoracic Surgeons database of 80,606 patients who underwent coronary artery bypass grafting in 1993. The results of traditional logistic regression and Bayesian analysis were compared with single-layer (no hidden layer), two-layer (one hidden layer), and three-layer (two hidden layer) MLP neural networks. These networks were trained using stochastic gradient descent with early stopping. All prediction models used the same variables and were evaluated by training on 40,480 patients and cross-validation testing on a separate group of 40,126 patients. Techniques were also developed to calculate effective odds ratios for MLP networks and to generate confidence intervals for MLP risk predictions using an auxiliary "confidence MLP." RESULTS: Receiver operating characteristic curve areas for predicting mortality were approximately 76% for all classifiers, including neural networks. Calibration (accuracy of posterior probability prediction) was slightly better with a two-member committee classifier that averaged the outputs of a MLP network and a logistic regression model. Unlike the individual methods, the committee classifier did not overestimate or underestimate risk for high-risk patients. CONCLUSIONS: A committee classifier combining the best neural network and logistic regression provided the best model calibration, but the receiver operating characteristic curve area was only 76% irrespective of which predictive model was used.

Bayes Theorem↗

Genetic neural network modeling of the selective inhibition of the intermediate-conductance Ca2+ -activated K+ channel by some triarylmethanes using topological charge indexes descriptors.

Selective inhibition of the intermediate-conductance Ca(2+)-activated K(+ )channel (IK (Ca)) by some clotrimazole analogs has been successfully modeled using topological charge indexes (TCI) and genetic neural networks (GNNs). A neural network monitoring scheme evidenced a highly non-linear dependence between the IK (Ca) blocking activity and TCI descriptors. Suitable subsets of descriptors were selected by means of genetic algorithm. Bayesian regularization was implemented in the network training function with the aim of assuring good generalization qualities to the predictors. GNNs were able to yield a reliable predictor that explained about 97% data variance with good predictive ability. On the contrary, the best multivariate linear equation with descriptors selected by linear genetic search, only explained about 60%. In spite of when using the descriptors from the linear equations to train neural networks yielded higher fitted models, such networks were very unstable and had relative low predictive ability. However, the best GNN BRANN 2 had a Q ( 2 ) of LOO of cross-validation equal to 0.901 and at the same time exhibited outstanding stability when calculating 80 randomly constructed training/test sets partitions. Our model suggested that structural fragments of size three and seven have relevant influence on the inhibitory potency of the studied IK (Ca) channel blockers. Furthermore, inhibitors were well distributed regarding its activity levels in a Kohonen self-organizing map (KSOM) built using the inputs of the best neural network predictor.

Algorithms↗

Recognizing complex, asymmetric functional sites in protein structures using a Bayesian scoring function.

The increase in known three-dimensional protein structures enables us to build statistical profiles of important functional sites in protein molecules. These profiles can then be used to recognize sites in large-scale automated annotations of new protein structures. We report an improved FEATURE system which recognizes functional sites in protein structures. FEATURE defines multi-level physico-chemical properties and recognizes sites based on the spatial distribution of these properties in the sites' microenvironments. It uses a Bayesian scoring function to compare a query region with the statistical profile built from known examples of sites and control nonsites. We have previously shown that FEATURE can accurately recognize calcium-binding sites and have reported interesting results scanning for calcium-binding sites in the entire Protein Data Bank. Here we report the ability of the improved FEATURE to characterize and recognize geometrically complex and asymmetric sites such as ATP-binding sites and disulfide bond-forming sites. FEATURE does not rely on conserved residues or conserved residue geometry of the sites. We also demonstrate that, in the absence of a statistical profile of the sites, FEATURE can use an artificially constructed profile based on a priori knowledge to recognize the sites in new structures, using redoxin active sites as an example.

Adenosine Triphosphate↗

High-dimensional image registration using symmetric priors.

This paper is about warping a brain image from one subject (the object image) so that it matches another (the template image). A high-dimensional model is used, whereby a finite element approach is employed to estimate translations at the location of each voxel in the template image. Bayesian statistics are used to obtain a maximum a posteriori (MAP) estimate of the deformation field. The validity of any registration method is largely based upon the constraints or, in this instance, priors incorporated into the model describing the transformations. In this approach we assume that the priors should have some form of symmetry, in that priors describing the probability distribution of the deformations should be identical to those for the inverses (i.e., warping brain A to brain B should not be different probabilistically from warping B to A). The fundamental assumption is that the probability of stretching a voxel by a factor of n is considered to be the same as the probability of shrinking n voxels by a factor of n(-1). In the Bayesian framework adopted here, the priors are assumed to have a Gibbs form, where the Gibbs potential is a penalty function that embodies this symmetry. The penalty function of choice is based upon the singular values of the Jacobian having a lognormal distribution. This enforces a continuous one-to-one mapping. A gradient descent algorithm is presented that incorporates the above priors in order to obtain a MAP estimate of the deformations. We demonstrate this approach for the two-dimensional case, but the principles can be extended to three dimensions. A number of examples are given to demonstrate how the method works.

Bayes Theorem↗

Application of template matching technique to particle detection in electron micrographs.

Template matching together with the comprehensive theory of image formation in electron microscope provides an optimal (in Bayesian sense) tool for solving one of the outstanding problems in single particle analysis, i.e., automatic selection of particle views from noisy micrograph fields. The method is based on the assumption that the reference three-dimensional structure is known and that the relevant parameters of the model of the image formation process can be estimated. In the first stage of the procedure, a set of possible particle views is generated using the available reference structure. The template images are constructed as linear combinations of available particle views using a clustering technique. Next, the micrograph noise characteristic is established using an automated contrast transfer function (CTF) estimation procedure. Finally, the CTF parameters calculated are used to construct a matched filter and correlation functions corresponding to the available template images are calculated. In order to alleviate the problem of the biased caused by varying image formation conditions, a decision making strategy based on the predicted distribution of correlation coefficients is proposed. It is demonstrated that due to the inclusion of CTF considerations, the template matching method performed very well in a broad range of microscopy conditions.

Algorithms↗