Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Collimator optimization for lesion detection incorporating prior information about lesion size.

A Bayesian estimator has been developed as a paradigm for human observer performance in detecting lesions of unknown size in a uniform noisy background. The Bayesian observer used knowledge of the range of possible lesion sizes as a prior; its predictions agreed well with the results of a six-observer perceptual study. The average human response to changes in collimator resolution, as measured by the detectability index, dA, was tracked by the Bayesian detector's signal-to-noise ratio (SNR) somewhat better than by two other estimation models based, respectively, on lesser and greater degrees of lesion size uncertainty. As the range of possible lesion sizes increased, the Bayesian detector's SNR decreased and the optimal collimator resolution shifted towards better resolution. An analytic approximation for the variance of lesion activity estimates (which included the same prior) was shown to predict the variance of the Bayesian estimator over a wide range of collimator resolution values. Because the bias of the Bayesian estimator was small (< 1%), the analytic variance estimate permitted a rapid and convenient prediction of the Bayesian detection SNR. This calculation was then used to optimize the geometric parameters of a two-layer tungsten collimator being constructed from crossed grids for a new imaging detector. A Monte Carlo program was first run to estimate all contributions to the radial point-spread function for collimators of differing tungsten contents and spatial resolution values, imaging 140-keV photons emitted from the center of a 15-cm-diameter, water-filled attenuator. The optimal collimator design for detecting lesions with unknown diameters in the range 2.5-7.5 mm yielded a system resolution of approximately 8.5-mm FWHM, a geometric collimator efficiency of 1.21 x 10(-4), and a single-septum penetration probability of 1%.

Bayes Theorem↗

Bayesian identification of differentially expressed isoforms using a novel joint model of RNA-seq data.

We develop a Bayesian approach, BayesIso, to identify differentially expressed isoforms from RNA-seq data. The approach features a novel joint model of the sample variability and the deferential state of isoforms. Specifically, the within-sample variability and the between-sample variability of each isoform are modeled by a Poisson-Lognormal model and a Gamma-Gamma model, respectively. Using a Bayesian framework, the differential state of each isoform and the model parameters are jointly estimated by a Markov Chain Monte Carlo (MCMC) method. Extensive studies using simulation and real data demonstrate that BayesIso can effectively detect isoforms of less differentially expressed and differential transcripts for genes with multiple isoforms. We applied the approach to breast cancer RNA-seq data and uncovered a unique set of isoforms that form key pathways associated with breast cancer recurrence. First, PI3K/AKT/mTOR signaling and PTEN signaling pathways are identified as being involved in breast cancer development. Further integrated with protein-protein interaction data, pathways of Jak-STAT, mTOR, MAPK and Wnt signaling are revealed in association with breast cancer recurrence. Finally, several pathways are activated in the early recurrence of breast cancer. In tumors that occur early, members of pathways of cellular metabolism and cell cycle (such as CD36 and TOP2A) are upregulated, while immune response genes such as NFATC1 are downregulated.

Humans↗

Segregation of visceral and somatosensory afferents: an fMRI and cytoarchitectonic mapping study.

Ano-rectal stimulation provides an important model for the processing of somatosensory and visceral sensations in the human nervous system. In spite of their anatomical proximity, the anal canal is innervated by somatosensory afferents whereas the rectum is innervated by the visceral nervous system. In a functional magnetic resonance (fMRI) experiment, we examined the cerebral responses to pneumatic balloon distension of these two structures to test whether somatosensory and visceral stimulation elicited distinct brain activations in spite of their spinal convergence. The specificity of the identified activations was analyzed by Bayesian mixed effects modeling. Activations in the parietal operculum were also compared to the location of cytoarchitectonically defined areas OP 1-4, which are part of the secondary somatosensory cortex (SII), to analyze whether the SII region was activated by anal and/or rectal stimulation. The lowest segregation between visceral and somatosensory stimuli was in the insular cortex, which supports the interpretation of the insula as an integrative region, receiving input from different sensory modalities. The most distinct segregation was found in the fronto-parietal operculum. Here the activations following anal and rectal stimulation were not only functionally but also anatomically distinct. Anal sensations were processed similar to other somatosensory stimuli in the SII cortex (area OP 4). Rectal afferents on the other hand were not processed in SII. Rather, they evoked activation at a more anterior location on the precentral operculum. These results demonstrate a functionally and anatomically distinct processing of somatosensory and visceral afferents in the human cerebral cortex.

Adult↗

Multiple-trait Gibbs sampler for animal models: flexible programs for Bayesian and likelihood-based (co)variance component inference.

A set of FORTRAN programs to implement a multiple-trait Gibbs sampling algorithm for (co)variance component inference in animal models (MTGSAM) was developed. The MTGSAM programs are available to the public. The programs support models with correlated genetic effects and arbitrary numbers of covariates, fixed effects, and independent random effects for each trait. Any combination of missing traits is allowed. The programs were used to estimate variance components for 50 replicates of simulated data. Each replicate consisted of 50 animals of each sex in each of four generations, for 400 animals in each replicate for two traits. For MTGSAM, informative prior distributions for variance components were inverted Wishart random variables with 10 df and means equal to the simulation parameters. A total of 15,000 Gibbs sampling rounds were completed for each replicate, with 2,000 rounds discarded for burn-in. For multiple-trait derivative free restricted maximum likelihood (MTDFREML), starting values for the variance components were the simulation parameters. Averages of posterior mean of variance components estimated using MTGSAM with informative and flat prior distributions for variance components and REML estimates obtained using MTDFREML indicated that all three methods were empirically unbiased. Correlations between estimates from MTGSAM using flat priors and MTDFREML all exceeded .99.

Algorithms↗

Maximum-likelihood versus maximum a posteriori parameter estimation of physiological system models: the C-peptide impulse response case study.

Maximum-likelihood (ML), also given its connection to least-squares (LS), is widely adopted in parameter estimation of physiological system models, i.e., assigning numerical values to the unknown model parameters from the experimental data. A more sophisticated but less used approach is maximum a posteriori (MAP) estimation. Conceptually, while ML adopts a Fisherian approach, i.e., only experimental measurements are supplied to the estimator, MAP estimation is a Bayesian approach, i.e., a priori available statistical information on the unknown parameters is also exploited for their estimation. In this paper, after a brief review of the theory behind ML and MAP estimators, we compare their performance in the solution of a case study concerning the determination of the parameters of a sum of exponential model which describes the impulse response of C-peptide (CP), a key substance for reconstructing insulin secretion. The results show that MAP estimation always leads to parameter estimates with a precision (sometimes significantly) higher than that obtained through ML, at the cost of only a slightly worse fit. Thus, a three exponential model can be adopted to describe the CP impulse response model in place of the two exponential model usually identified in the literature by the ML/LS approach. Simulated case studies are also reported to evidence the importance of taking into account a priori information in a data poor situation, e.g., when a few or too noisy measurements are available. In conclusion, our results show that, when a priori information on the unknown model parameters is available, Bayes estimation can be of relevant interest, since it can significantly improve the precision of parameter estimates with respect to Fisher estimation. This may also allow the adoption of more complex models than those determinable by a Fisherian approach.

Bayes Theorem↗

Unsupervised learning with independent component analysis can identify patterns of glaucomatous visual field defects.

PURPOSE: We previously reported the use of clustering by unsupervised learning with machine learning classifiers to segment clusters of patterns in standard automated perimetry (SAP) for glaucoma. In this study, the process of unsupervised learning by independent component analysis decomposed SAP field patterns into axes, and the information represented by these axes was evaluated. METHODS: SAP fields were obtained with the Humphrey Visual Field Analyzer on 189 normal eyes and 156 eyes with glaucomatous optic neuropathy (GON) determined by masked review with stereoscopic optic disc photos. The variational Bayesian independent component analysis mixture model (vB-ICA-mm) partitioned the SAP fields into the most informative number of clusters. Simultaneously, it learned an optimal number of maximally independent axes for each cluster. RESULTS: The most informative number of clusters was two. vB-ICA-mm placed 68.6% of the SAP fields from eyes with GON in a cluster labeled G and 98.4% of the fields from eyes with normal optic discs in a cluster labeled N. Cluster G optimally contained six axes. Post hoc analysis of patterns generated at -1 SD and +2 SD from the cluster G mean on the six axes revealed defects similar to those identified by experts as indicative of glaucoma. SAP fields associated with an axis showed increasing severity as they were located farther in the positive direction from the cluster G mean. CONCLUSIONS: vB-ICA-mm represented the SAP fields with patterns that were meaningful for glaucoma experts. This process also captured severity in the patterns uncovered. These findings should validate vB-ICA-mm as a data mining technique for new and unfamiliar complex tests.

Artificial Intelligence↗

Comparison of rheumatological diagnoses by a Bayesian program and by physicians.

A Bayesian decision support system was developed for the diagnosis of rheumatic disorders. Knowledge in this system is represented as evidential weights of findings. Simple weights were calculated as the logarithm of likelihood ratios on the basis of 1,000 consecutive patients from a rheumatological clinic. The effect of various methods to improve performance of the system by modification of the weights was studied. Three methods had a mathematical basis; a fourth consisted of weights adapted by a human expert, which allowed inclusion of diagnostic rules such as defined in widely accepted criteria sets. The system's performance was measured in a test population of 570 different cases from the same clinic and compared with predictions of diagnostic outcome made by rheumatologists. The weights from a human expert gave optimal results (sensitivity 65% and specificity 96%), that were close to the physicians' predictions (sensitivity 64% and specificity 98%). The methods to measure the performance of the various models used in this study emphasize sensitivity, specificity and the use of receiver operating characteristics.

Adolescent↗

Bayesian inference in a hidden stochastic two-compartment model for feline hematopoiesis.

In this paper, we describe a hidden two-compartment stochastic process used to model the kinetics of feline hematopoietic stem cells (HSCs) in continuous time. Because of the experimental design and data collection scheme, the inferential task presents numerous challenges. While the hematopoietic process evolves in continuous time, the observations are collected only at discrete irregular times and are a probabilistic function of the state of the process. In addition, the animals go through an experimental procedure such that their reserve of HSCs is severely depleted at the start of the observation period. This impedes any approximation of the hematopoietic process with a continuous state-space process (normal approximation of the transition probabilities would be inaccurate when the state of the process, i.e. the number of stem cells, is small). We implement a Markov chain Monte Carlo algorithm that allows us to estimate the posterior distribution of the parameters of the hematopoietic process while maintaining its state-space discrete (i.e. without using any approximation). We show the performance of the algorithm on simulated data. Finally, we apply the algorithm to data on multiple experimental cats and provide estimates of the rates of the fates of feline HSCs. The obtained estimates are in agreement with the estimates obtained with different methods published in the medical literature. However, the proposed approach makes a more efficient use of the data and hence the parameter estimates are much more accurate than the one obtained with the methods previously proposed.

Algorithms↗

Bayesian fluorescence in situ hybridisation signal classification.

Previous research has indicated the significance of accurate classification of fluorescence in situ hybridisation (FISH) signals for the detection of genetic abnormalities. Based on well-discriminating features and a trainable neural network (NN) classifier, a previous system enabled highly-accurate classification of valid signals and artefacts of two fluorophores. However, since this system employed several features that are considered independent, the naive Bayesian classifier (NBC) is suggested here as an alternative to the NN. The NBC independence assumption permits the decomposition of the high-dimensional likelihood of the model for the data into a product of one-dimensional probability densities. The naive independence assumption together with the Bayesian methodology allow the NBC to predict a posteriori probabilities of class membership using estimated class-conditional densities in a close and simple form. Since the probability densities are the only parameters of the NBC, the misclassification rate of the model is determined exclusively by the quality of density estimation. Densities are evaluated by three methods: single Gaussian estimation (SGE; parametric method), Gaussian mixture model assuming spherical covariance matrices (GMM; semi-parametric method) and kernel density estimation (KDE; non-parametric method). For low-dimensional densities, the GMM generally outperforms the KDE that tends to overfit the training set at the cost of reduced generalisation capability. But, it is the GMM that loses some accuracy when modelling higher-dimensional densities due to the violation of the assumption of spherical covariance matrices when dependent features are added to the set. Compared with these two methods, the SGE and NN provide inferior and superior performance, respectively. However, the NBC avoids the intensive training and optimisation required for the NN, demanding extensive resources and experimentation. Therefore, when supporting these two classifiers, the system enables a trade-off between the NN performance and NBC simplicity of implementation.

Algorithms↗

Protein structure elucidation from NMR proton densities.

The NMR-generated foc proton density affords a template to which the molecule has to be fitted to derive the structure. Here we present a computational protocol that achieves this goal. H(N) atoms are readily recognizable from (1)H/(2)H exchange or (1)H/(15)N heteronuclear single quantum correlation (HSQC) experiments. The primary structure is threaded through the unassigned foc by leapfrogging along peptidyl amide H(N)s and the connected H(alpha)s. Via a Bayesian approach, the probabilities of the sequential connectivity hypotheses are inferred from likelihoods of H(N)/H(N), H(N)/H(alpha), and H(alpha)/H(alpha) interatomic distances as well as (1)H NMR chemical shifts, both derived from public databases. Once the polypeptide sequence is identified, directionality becomes established, and the foc N and C termini are recognized. After a similar procedure, side chain H atoms are found, including discriminated cis/trans proline loci. The folded structure then is derived via a direct molecular dynamics embedding into mirror image-related representations of the foc and selected according to a lowest energy criterion. The method was applied to foc densities calculated for two protein domains, col 2 and kringle 2. The obtained structures are within 1.0-1.5 A (backbone heavy atoms) and 1.5-2.0 A (all heavy atoms) rms deviations from reported x-ray and/or NMR structures.

Computer Simulation↗

Evaluation of bayesian forecasting for individualized gentamicin dosage in infants weighing 1000 g or less.

We evaluated the use of Bayesian forecasting for gentamicin therapy in outborn infants weighing 1000 g or less irrespective of postnatal age. Dosages were individualized using a computer program, guided by early serum gentamicin assays after a loading dose and a database of population kinetics. Steady-state gentamicin levels achieved were compared with those from a regimen based on guidelines. A total of 26 gentamicin courses were individualized in 19 infants of 22 to 33 weeks' gestation, weighing 500 to 1000 g at 1 to 41 days of age. All steady-state trough levels were between 1 and 2.4 mg/L; peak levels were between 4.4 and 9.3 mg/L. The 95% confidence intervals were in almost identical ranges. The prevalence of toxic and suboptimal trough levels was less when compared with that of 23 gentamicin courses based on guidelines in 17 control infants. We conclude that early individualized gentamicin dosage over a range of postnatal age is a practical alternative and serum level distributions appear superior.

Bayes Theorem↗

Geostatistical analysis of disease data: estimation of cancer mortality risk from empirical frequencies using Poisson kriging.

BACKGROUND: Cancer mortality maps are used by public health officials to identify areas of excess and to guide surveillance and control activities. Quality of decision-making thus relies on an accurate quantification of risks from observed rates which can be very unreliable when computed from sparsely populated geographical units or recorded for minority populations. This paper presents a geostatistical methodology that accounts for spatially varying population sizes and spatial patterns in the processing of cancer mortality data. Simulation studies are conducted to compare the performances of Poisson kriging to a few simple smoothers (i.e. population-weighted estimators and empirical Bayes smoothers) under different scenarios for the disease frequency, the population size, and the spatial pattern of risk. A public-domain executable with example datasets is provided. RESULTS: The analysis of age-adjusted mortality rates for breast and cervix cancers illustrated some key features of commonly used smoothing techniques. Because of the small weight assigned to the rate observed over the entity being smoothed (kernel weight), the population-weighted average leads to risk maps that show little variability. Other techniques assign larger and similar kernel weights but they use a different piece of auxiliary information in the prediction: global or local means for global or local empirical Bayes smoothers, and spatial combination of surrounding rates for the geostatistical estimator. Simulation studies indicated that Poisson kriging outperforms other approaches for most scenarios, with a clear benefit when the risk values are spatially correlated. Global empirical Bayes smoothers provide more accurate predictions under the least frequent scenario of spatially random risk. CONCLUSION: The approach presented in this paper enables researchers to incorporate the pattern of spatial dependence of mortality rates into the mapping of risk values and the quantification of the associated uncertainty, while being easier to implement than a full Bayesian model. The availability of a public-domain executable makes the geostatistical analysis of health data, and its comparison to traditional smoothers, more accessible to common users. In future papers this methodology will be generalized to the simulation of the spatial distribution of risk values and the propagation of the uncertainty attached to predicted risks in local cluster analysis.

Journal Article↗

Linkage analysis of quantitative trait loci in multiple line crosses.

Simple line crosses, for example, backcross and F2, are commonly used in mapping quantitative trait loci (QTL). However, these simple crosses are rarely used alone in commercial plant breeding; rather, crosses involving multiple inbred lines or several simple crosses but connected by shared inbred lines may be common in plant breeding. Mapping QTL using crosses of multiple lines is more relevant to plant breeding. Unfortunately, current statistical methods and computer programs of QTL mapping are all designed for simple line crosses or multiple line crosses but under a regular mating system. It is not straightforward to extend the existing methods to handle multiple line crosses under irregular and complicated mating designs. The major hurdle comes from irregular inbreeding, multiple generations, and multiple alleles. In this study, we develop a Bayesian method implemented via the Markov chain Monte Carlo (MCMC) algorithm for mapping QTL using complicated multiple line crosses. With the MCMC algorithm, we are able to draw a complete path of the gene flow from founder alleles to their descendents via a recursive process. This has greatly simplified the problem caused by irregular mating and inbreeding in the mapping population. Adopting the reversible jump MCMC algorithm, we are able to simultaneously search for multiple QTL along the genome. We can even infer the posterior distribution of the number of QTL, one of the most important parameters in QTL study. Application of the new MCMC based QTL mapping procedure is demonstrated using two different mating designs. Design I involves two inbred lines and their derived F1, F2, and BC populations. Design II is a half-diallel cross involving three inbred lines. The two designs appear different, but can be handled with the same robust computer program.

Algorithms↗

Detecting coevolving amino acid sites using Bayesian mutational mapping.

MOTIVATION: The evolution of protein sequences is constrained by complex interactions between amino acid residues. Because harmful substitutions may be compensated for by other substitutions at neighboring sites, residues can coevolve. We describe a Bayesian phylogenetic approach to the detection of coevolving residues in protein families. This method, Bayesian mutational mapping (BMM), assigns mutations to the branches of the evolutionary tree stochastically, and then test statistics are calculated to determine whether a coevolutionary signal exists in the mapping. Posterior predictive P-values provide an estimate of significance, and specificity is maintained by integrating over uncertainty in the estimation of the tree topology, branch lengths and substitution rates. A coevolutionary Markov model for codon substitution is also described, and this model is used as the basis of several test statistics. RESULTS: Results on simulated coevolutionary data indicate that the BMM method can successfully detect nearly all coevolving sites when the model has been correctly specified, and that non-parametric statistics such as mutual information are generally less powerful than parametric statistics. On a dataset of eukaryotic proteins from the phosphoglycerate kinase (PGK) family, interdomain site contacts yield a significantly greater coevolutionary signal than interdomain non-contacts, an indication that the method provides information about interacting sites. Failure to account for the heterogeneity in rates across sites in PGK resulted in a less discriminating test, yielding a marked increase in the number of reported positives at both contact and non-contact sites. SUPPLEMENTARY INFORMATION: http://www.dimmic.net/supplement/

Amino Acids↗

Integration of association statistics over genomic regions using Bayesian adaptive regression splines.

In the search for genetic determinants of complex disease, two approaches to association analysis are most often employed, testing single loci or testing a small group of loci jointly via haplotypes for their relationship to disease status. It is still debatable which of these approaches is more favourable, and under what conditions. The former has the advantage of simplicity but suffers severely when alleles at the tested loci are not in linkage disequilibrium (LD) with liability alleles; the latter should capture more of the signal encoded in LD, but is far from simple. The complexity of haplotype analysis could be especially troublesome for association scans over large genomic regions, which, in fact, is becoming the standard design. For these reasons, the authors have been evaluating statistical methods that bridge the gap between single-locus and haplotype-based tests. In this article, they present one such method, which uses non-parametric regression techniques embodied by Bayesian adaptive regression splines (BARS). For a set of markers falling within a common genomic region and a corresponding set of single-locus association statistics, the BARS procedure integrates these results into a single test by examining the class of smooth curves consistent with the data. The non-parametric BARS procedure generally finds no signal when no liability allele exists in the tested region (ie it achieves the specified size of the test) and it is sensitive enough to pick up signals when a liability allele is present. The BARS procedure provides a robust and potentially powerful alternative to classical tests of association, diminishes the multiple testing problem inherent in those tests and can be applied to a wide range of data types, including genotype frequencies estimated from pooled samples.

Algorithms↗

Donuts, scratches and blanks: robust model-based segmentation of microarray images.

MOTIVATION: Inner holes, artifacts and blank spots are common in microarray images, but current image analysis methods do not pay them enough attention. We propose a new robust model-based method for processing microarray images so as to estimate foreground and background intensities. The method starts with a very simple but effective automatic gridding method, and then proceeds in two steps. The first step applies model-based clustering to the distribution of pixel intensities, using the Bayesian Information Criterion (BIC) to choose the number of groups up to a maximum of three. The second step is spatial, finding the large spatially connected components in each cluster of pixels. The method thus combines the strengths of the histogram-based and spatial approaches. It deals effectively with inner holes in spots and with artifacts. It also provides a formal inferential basis for deciding when the spot is blank, namely when the BIC favors one group over two or three. RESULTS: We apply our methods for gridding and segmentation to cDNA microarray images from an HIV infection experiment. In these experiments, our method had better stability across replicates than a fixed-circle segmentation method or the seeded region growing method in the SPOT software, without introducing noticeable bias when estimating the intensities of differentially expressed genes. AVAILABILITY: spotSegmentation, an R language package implementing both the gridding and segmentation methods is available through the Bioconductor project (http://www.bioconductor.org). The segmentation method requires the contributed R package MCLUST for model-based clustering (http://cran.us.r-project.org). CONTACT: fraley@stat.washington.edu.

Algorithms↗

Estimation of gentamicin clearance and volume of distribution in neonates and young children.

Gentamicin therapy should be guided by serum level monitoring in all age groups, dosage adjustments depending on age related changes in pharmacokinetics. Population data analysed from two centres (43 infants from Glasgow and 100 infants and children from Manchester) by the computer program NONMEM showed that volume of distribution was related to body weight by a proportionality factor that decreased from the region of 0.41-0.46 l/kg in children less than 3 months to 0.25-0.32 l/kg in older children, a value which merges with that accepted for adults (0.25 l/kg). In both young and older children, clearance was also found to be dependent on body weight. Renal function (creatinine concentrations) provided no further explanatory power. When these results were used prospectively to forecast gentamicin concentrations with a Bayesian kinetic parameter estimation program, trough concentrations were more precisely predicted than peaks when a single concentration measurement was used. In clinical practice, however, two concentration measurements are usually routinely available and these should lead to greater precision of both peak and trough predictions. These results have been incorporated into a simple nomogram which can be used to determine a dose of gentamicin which will achieve target peak concentrations in infants, assuming that troughs should not exceed 2 micrograms/ml.

Age Factors↗

The incorporation of prior genomic information does not necessarily improve the performance of Bayesian linkage methods: an example involving sex-specific recombination and the two-point PPL.

OBJECTIVE: We continue statistical development of the posterior probability of linkage (PPL). We present a two-point PPL allowing for unequal male and female recombination fractions, thetaM and thetaF, and consider alternative priors on thetaM, thetaF. METHODS: We compare the sex-averaged PPL (PPLSA), assuming thetaM = thetaF, to the sex-specific PPL (PPLSS) in (thetaM, thetaF), in a series of simulations; we also compute the PPLSS using alternative priors on (thetaM, thetaF). RESULTS: The PPLSS based on a prior that ignores prior genomic information on sex specific recombination rates performs essentially identically to the PPLSA, even in the presence of large thetaM, thetaF differences. Moreover, adaptively skewing the prior, to incorporate (correct) genomic information on thetaM, thetaF differences, actually worsens performance of the PPLSS. We demonstrate that this has little to do with the PPLSS per se, but is rather due to extremely high levels of variability in the location of the maximum likelihood estimates of (thetaM, thetaF) in realistic data sets. CONCLUSIONS: Incorporating (correct) prior genomic information is not always helpful. We recommend that the PPLSA be used as the standard form of the PPL regardless of the sex-specific recombination rates in the region of the marker in question.

Bayes Theorem↗