Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

Design and analysis of admixture mapping studies.

Admixture between populations originating on different continents can be exploited to detect disease susceptibility loci at which risk alleles are distributed differentially between these populations. We first examine the statistical power and mapping resolution of this approach in the limiting situation in which gamete admixture and locus ancestry are measured without uncertainty. We show that, for a rare disease, the most efficient design is to study affected individuals only. In a typical African American population (two-way admixture proportions 0.8/0.2, ancestry crossover rate 2 per 100 cM), a study of 800 affected individuals has 90% power to detect at P values <10(-5) a locus that generates a risk ratio of 2 between populations, with an expected mapping resolution (size of 95% confidence region for the position of the locus) of 4 cM. In practice, to infer locus ancestry from marker data requires Bayesian computationally intensive methods, as implemented in the program ADMIXMAP. Affected-only study designs require strong prior information on the frequencies of each allele given locus ancestry. We show how data from unadmixed and admixed populations can be combined to estimate these ancestry-specific allele frequencies within the admixed population under study, allowing for variation between allele frequencies in unadmixed and admixed populations. Using simulated data based on the genetic structure of the African American population, we show that 60% of information can be extracted in a test for linkage using markers with an ancestry information content of 36% at 3-cM spacing. As in classic linkage studies, the most efficient strategy is to use markers at a moderate density for an initial genome search and then to saturate regions of putative linkage with additional markers, to extract nearly all information about locus ancestry.

Black People↗

Determination of strongly overlapping signaling activity from microarray data.

BACKGROUND: As numerous diseases involve errors in signal transduction, modern therapeutics often target proteins involved in cellular signaling. Interpretation of the activity of signaling pathways during disease development or therapeutic intervention would assist in drug development, design of therapy, and target identification. Microarrays provide a global measure of cellular response, however linking these responses to signaling pathways requires an analytic approach tuned to the underlying biology. An ongoing issue in pattern recognition in microarrays has been how to determine the number of patterns (or clusters) to use for data interpretation, and this is a critical issue as measures of statistical significance in gene ontology or pathways rely on proper separation of genes into groups. RESULTS: Here we introduce a method relying on gene annotation coupled to decompositional analysis of global gene expression data that allows us to estimate specific activity on strongly coupled signaling pathways and, in some cases, activity of specific signaling proteins. We demonstrate the technique using the Rosetta yeast deletion mutant data set, decompositional analysis by Bayesian Decomposition, and annotation analysis using ClutrFree. We determined from measurements of gene persistence in patterns across multiple potential dimensionalities that 15 basis vectors provides the correct dimensionality for interpreting the data. Using gene ontology and data on gene regulation in the Saccharomyces Genome Database, we identified the transcriptional signatures of several cellular processes in yeast, including cell wall creation, ribosomal disruption, chemical blocking of protein synthesis, and, critically, individual signatures of the strongly coupled mating and filamentation pathways. CONCLUSION: This works demonstrates that microarray data can provide downstream indicators of pathway activity either through use of gene ontology or transcription factor databases. This can be used to investigate the specificity and success of targeted therapeutics as well as to elucidate signaling activity in normal and disease processes.

Algorithms↗

A spatial statistical model for landscape genetics.

Landscape genetics is a new discipline that aims to provide information on how landscape and environmental features influence population genetic structure. The first key step of landscape genetics is the spatial detection and location of genetic discontinuities between populations. However, efficient methods for achieving this task are lacking. In this article, we first clarify what is conceptually involved in the spatial modeling of genetic data. Then we describe a Bayesian model implemented in a Markov chain Monte Carlo scheme that allows inference of the location of such genetic discontinuities from individual geo-referenced multilocus genotypes, without a priori knowledge on populational units and limits. In this method, the global set of sampled individuals is modeled as a spatial mixture of panmictic populations, and the spatial organization of populations is modeled through the colored Voronoi tessellation. In addition to spatially locating genetic discontinuities, the method quantifies the amount of spatial dependence in the data set, estimates the number of populations in the studied area, assigns individuals to their population of origin, and detects individual migrants between populations, while taking into account uncertainty on the location of sampled individuals. The performance of the method is evaluated through the analysis of simulated data sets. Results show good performances for standard data sets (e.g., 100 individuals genotyped at 10 loci with 10 alleles per locus), with high but also low levels of population differentiation (e.g., FST<0.05). The method is then applied to a set of 88 individuals of wolverines (Gulo gulo) sampled in the northwestern United States and genotyped at 10 microsatellites.

Animals↗

Active and dynamic information fusion for facial expression understanding from image sequences.

This paper explores the use of multisensory information fusion technique with Dynamic Bayesian networks (DBNs) for modeling and understanding the temporal behaviors of facial expressions in image sequences. Our facial feature detection and tracking based on active IR illumination provides reliable visual information under variable lighting and head motion. Our approach to facial expression recognition lies in the proposed dynamic and probabilistic framework based on combining DBNs with Ekman's Facial Action Coding System (FACS) for systematically modeling the dynamic and stochastic behaviors of spontaneous facial expressions. The framework not only provides a coherent and unified hierarchical probabilistic framework to represent spatial and temporal information related to facial expressions, but also allows us to actively select the most informative visual cues from the available information sources to minimize the ambiguity in recognition. The recognition of facial expressions is accomplished by fusing not only from the current visual observations, but also from the previous visual evidences. Consequently, the recognition becomes more robust and accurate through explicitly modeling temporal behavior of facial expression. In this paper, we present the theoretical foundation underlying the proposed probabilistic and dynamic framework for facial expression modeling and understanding. Experimental results demonstrate that our approach can accurately and robustly recognize spontaneous facial expressions from an image sequence under different conditions.

Algorithms↗

Bayesian probabilistic approach for predicting backbone structures in terms of protein blocks.

By using an unsupervised cluster analyzer, we have identified a local structural alphabet composed of 16 folding patterns of five consecutive C(alpha) ("protein blocks"). The dependence that exists between successive blocks is explicitly taken into account. A Bayesian approach based on the relation protein block-amino acid propensity is used for prediction and leads to a success rate close to 35%. Sharing sequence windows associated with certain blocks into "sequence families" improves the prediction accuracy by 6%. This prediction accuracy exceeds 75% when keeping the first four predicted protein blocks at each site of the protein. In addition, two different strategies are proposed: the first one defines the number of protein blocks in each site needed for respecting a user-fixed prediction accuracy, and alternatively, the second one defines the different protein sites to be predicted with a user-fixed number of blocks and a chosen accuracy. This last strategy applied to the ubiquitin conjugating enzyme (alpha/beta protein) shows that 91% of the sites may be predicted with a prediction accuracy larger than 77% considering only three blocks per site. The prediction strategies proposed improve our knowledge about sequence-structure dependence and should be very useful in ab initio protein modelling.

Artificial Intelligence↗

ARACNE: an algorithm for the reconstruction of gene regulatory networks in a mammalian cellular context.

BACKGROUND: Elucidating gene regulatory networks is crucial for understanding normal cell physiology and complex pathologic phenotypes. Existing computational methods for the genome-wide "reverse engineering" of such networks have been successful only for lower eukaryotes with simple genomes. Here we present ARACNE, a novel algorithm, using microarray expression profiles, specifically designed to scale up to the complexity of regulatory networks in mammalian cells, yet general enough to address a wider range of network deconvolution problems. This method uses an information theoretic approach to eliminate the majority of indirect interactions inferred by co-expression methods. RESULTS: We prove that ARACNE reconstructs the network exactly (asymptotically) if the effect of loops in the network topology is negligible, and we show that the algorithm works well in practice, even in the presence of numerous loops and complex topologies. We assess ARACNE's ability to reconstruct transcriptional regulatory networks using both a realistic synthetic dataset and a microarray dataset from human B cells. On synthetic datasets ARACNE achieves very low error rates and outperforms established methods, such as Relevance Networks and Bayesian Networks. Application to the deconvolution of genetic networks in human B cells demonstrates ARACNE's ability to infer validated transcriptional targets of the cMYC proto-oncogene. We also study the effects of misestimation of mutual information on network reconstruction, and show that algorithms based on mutual information ranking are more resilient to estimation errors. CONCLUSION: ARACNE shows promise in identifying direct transcriptional interactions in mammalian cellular networks, a problem that has challenged existing reverse engineering algorithms. This approach should enhance our ability to use microarray data to elucidate functional mechanisms that underlie cellular processes and to identify molecular targets of pharmacological compounds in mammalian cellular networks.

Algorithms↗

A general approach to mixed effects modeling of residual variances in generalized linear mixed models.

We propose a general Bayesian approach to heteroskedastic error modeling for generalized linear mixed models (GLMM) in which linked functions of conditional means and residual variances are specified as separate linear combinations of fixed and random effects. We focus on the linear mixed model (LMM) analysis of birth weight (BW) and the cumulative probit mixed model (CPMM) analysis of calving ease (CE). The deviance information criterion (DIC) was demonstrated to be useful in correctly choosing between homoskedastic and heteroskedastic error GLMM for both traits when data was generated according to a mixed model specification for both location parameters and residual variances. Heteroskedastic error LMM and CPMM were fitted, respectively, to BW and CE data on 8847 Italian Piemontese first parity dams in which residual variances were modeled as functions of fixed calf sex and random herd effects. The posterior mean residual variance for male calves was over 40% greater than that for female calves for both traits. Also, the posterior means of the standard deviation of the herd-specific variance ratios (relative to a unitary baseline) were estimated to be 0.60 +/- 0.09 for BW and 0.74 +/- 0.14 for CE. For both traits, the heteroskedastic error LMM and CPMM were chosen over their homoskedastic error counterparts based on DIC values.

Analysis of Variance↗

Using literature and data to learn Bayesian networks as clinical models of ovarian tumors.

Thanks to its increasing availability, electronic literature has become a potential source of information for the development of complex Bayesian networks (BN), when human expertise is missing or data is scarce or contains much noise. This opportunity raises the question of how to integrate information from free-text resources with statistical data in learning Bayesian networks. Firstly, we report on the collection of prior information resources in the ovarian cancer domain, which includes "kernel" annotations of the domain variables. We introduce methods based on the annotations and literature to derive informative pairwise dependency measures, which are derived from the statistical cooccurrence of the names of the variables, from the similarity of the "kernel" descriptions of the variables and from a combined method. We perform wide-scale evaluation of these text-based dependency scores against an expert reference and against data scores (the mutual information (MI) and a Bayesian score). Next, we transform the text-based dependency measures into informative text-based priors for Bayesian network structures. Finally, we report the benefit of such informative text-based priors on the performance of a Bayesian network for the classification of ovarian tumors from clinical data.

Artificial Intelligence↗

Bayesian models of episodic evolution support a late precambrian explosive diversification of the Metazoa.

Multicellular animals, or Metazoa, appear in the fossil records between 575 and 509 million years ago (MYA). At odds with paleontological evidence, molecular estimates of basal metazoan divergences have been consistently older than 700 MYA. However, those date estimates were based on the molecular clock hypothesis, which is almost always violated. To relax this hypothesis, we have implemented a Bayesian approach to describe the change of evolutionary rate over time. Analysis of 22 genes from the nuclear and the mitochondrial genomes under the molecular clock assumption produced old date estimates, similar to those from previous studies. However, by allowing rates to vary in time and by taking small species-sampling fractions into account, we obtained much younger estimates, broadly consistent with the fossil records. In particular, the date of protostome-deuterostome divergence was on average 582 +/- 112 MYA. These results were found to be robust to specification of the model of rate change. The clock assumption thus had a dramatic effect on date estimation. However, our results appeared sensitive to the prior model of cladogenesis, although the oldest estimates (791 +/- 246 MYA) were obtained under a suboptimal model. Bayes posterior estimates of evolutionary rates indicated at least one major burst of molecular evolution at the end of the Precambrian when protostomes and deuterostomes diverged. We stress the importance of assumptions about rates on date estimation and suggest that the large discrepancies between the molecular and fossil dates of metazoan divergences might partly be due to biases in molecular date estimation.

Algorithms↗

Relevance feedback using generalized Bayesian framework with region-based optimization learning.

This paper presents a generalized Bayesian framework for relevance feedback in content-based image retrieval. The proposed feedback technique is based on the Bayesian learning method and incorporates a time-varying user model into the formulation. We define the user model with two terms: a target query and a user conception. The target query is aimed to learn the common features from relevant images so as to specify the user's ideal query. The user conception is aimed to learn a parameter set to determine the time-varying matching criterion. Therefore, at each feedback step, the learning process updates not only the target distribution, but also the target query and the matching criterion. In addition, another objective of this paper is to conduct the relevance feedback on images represented in region level. We formulate the matching criterion using a weighting scheme and proposed a region clustering technique to determine the region correspondence between relevant images. With the proposed region clustering technique, we derive a representation in region level to characterize the target query. Experiments demonstrate that the proposed method combined with time-varying user model indeed achieves satisfactory results and our proposed region-based techniques further improve the retrieval accuracy.

Algorithms↗

Bayesian mapping of multiple quantitative trait loci from incomplete inbred line cross data.

A novel fine structure mapping method for quantitative traits is presented. It is based on Bayesian modeling and inference, treating the number of quantitative trait loci (QTLs) as an unobserved random variable and using ideas similar to composite interval mapping to account for the effects of QTLs in other chromosomes. The method is introduced for inbred lines and it can be applied also in situations involving frequent missing genotypes. We propose that two new probabilistic measures be used to summarize the results from the statistical analysis: (1) the (posterior) QTL intensity, for estimating the number of QTLs in a chromosome and for localizing them into some particular chromosomal regions, and (2) the locationwise (posterior) distributions of the phenotypic effects of the QTLs. Both these measures will be viewed as functions of the putative QTL locus, over the marker range in the linkage group. The method is tested and compared with standard interval and composite interval mapping techniques by using simulated backcross progeny data. It is implemented as a software package. Its initial version is freely available for research purposes under the name Multimapper at URL http://www.rni.helsinki.fi/mjs.

Bayes Theorem↗

Estimation of saturated pixel values in digital color imaging.

Pixel saturation, in which the incident light at a pixel causes one of the color channels of the camera sensor to respond at its maximum value, can produce undesirable artifacts in digital color images. We present a Bayesian algorithm that estimates what the saturated channel's value would have been in the absence of saturation. The algorithm uses the nonsaturated responses from the other color channels, together with a multivariate normal prior that captures the correlation in response across color channels. The prior may be estimated directly from the image data, since most image pixels are not saturated. Given the prior and the responses of the nonsaturated channels, the algorithm returns the optimal expected mean square estimate for the true response. Extensions of the algorithm to the case in which more than one channel is saturated are also discussed. Both simulations and examples with real images are presented to show that the algorithm is effective.

Algorithms↗

Fully Bayesian spatio-temporal modeling of FMRI data.

We present a fully Bayesian approach to modeling in functional magnetic resonance imaging (FMRI), incorporating spatio-temporal noise modeling and haemodynamic response function (HRF) modeling. A fully Bayesian approach allows for the uncertainties in the noise and signal modeling to be incorporated together to provide full posterior distributions of the HRF parameters. The noise modeling is achieved via a nonseparable space-time vector autoregressive process. Previous FMRI noise models have either been purely temporal, separable or modeling deterministic trends. The specific form of the noise process is determined using model selection techniques. Notably, this results in the need for a spatially nonstationary and temporally stationary spatial component. Within the same full model, we also investigate the variation of the HRF in different areas of the activation, and for different experimental stimuli. We propose a novel HRF model made up of half-cosines, which allows distinct combinations of parameters to represent characteristics of interest. In addition, to adaptively avoid over-fitting we propose the use of automatic relevance determination priors to force certain parameters in the model to zero with high precision if there is no evidence to support them in the data. We apply the model to three datasets and observe matter-type dependence of the spatial and temporal noise, and a negative correlation between activation height and HRF time to main peak (although we suggest that this apparent correlation may be due to a number of different effects).

Bayes Theorem↗

A new method for computing the multipoint posterior probability of linkage.

The posterior probability of linkage (PPL) is a Bayesian statistic which directly measures the probability of linkage between a trait locus and a marker (in the 2-point case) or a genomic region (in the multipoint case). It has several benefits, including ease of interpretation, the ability to incorporate prior genomic information, and a mathematically rigorous and robust procedure for accumulating linkage information across multiple heterogeneous datasets. To date, the majority of work on the PPL has focused on the development of the 2-point statistic, with only preliminary attempts at the development of an equivalent multipoint version. In this paper we present a new way of computing of the multipoint PPL. This new version imputes to each genomic point an estimate of the 2-point PPL we would have obtained from a fully informative marker giving similar evidence for linkage. This version, which we call the imputed PPL, is shown to be superior to previously developed versions.

Bayes Theorem↗

Estimation of mortality rates for disease simulation models using Bayesian evidence synthesis.

PURPOSE: The authors propose a Bayesian approach for estimating competing risks for inputs to disease simulation models. This approach is suggested when modeling a disease that causes a large proportion of all-cause mortality, particularly when mortality from the disease of interest and other-cause mortality are both affected by the same risk factor. METHODS: The authors demonstrate a Bayesian evidence synthesis by estimating other-cause mortality, stratified by smoking status, for use in a simulation model of lung cancer. National (US) survey data linked to death registries (National Health Interview Survey [NHIS]--Multiple Cause of Death files) were used to fit cause-specific hazard models for 3 causes of death (lung cancer, heart disease, and all other causes), controlling for age, sex, race, and smoking status. Synthesis of NHIS data with national vital statistics data on numbers and causes of deaths was performed in WinBUGS (version 1.4.1, MRC Biostatistics Unit, UK). Correction for inconsistencies between the NHIS and vital statistics data is described. A published cohort study was a source of prior information for smoking-related mortality. RESULTS: Marginal posterior densities of annual mortality rates for lung cancer and other-cause death (further divided into heart disease and all other causes), stratified by 5-year age interval, race (white and black), gender, and smoking status (current, former, never), were estimated, specific to a time period (1987-1995). Overall, black current smokers experienced the highest mortality rates. CONCLUSIONS: Bayesian evidence synthesis is an effective method for estimation of cause-specific mortality rates, stratified by demographic factors.

Adult↗

Population and individual minimal modeling of the frequently sampled insulin-modified intravenous glucose tolerance test.

Population approaches are more robust estimators of insulin sensitivity (SI) and glucose effectiveness (SG) with the minimal model of glucose kinetics during an intravenous glucose tolerance test (IVGTT). We assessed the performance of 3 population methods, iterative two-stage (ITS), Bayesian hierarchical Markov chain Monte Carlo (MCMC), and NONMEM first-order conditional estimation (FOCE) with interaction (NM), and made a comparison with the standard two-stage method (STS) employing the weighted nonlinear regression analysis. To evaluate accuracy of individual and population estimates, 40 simulated insulin-modified frequently sampled IVGTTs (IM-FSIVGTT) were derived from real IM-FSIVGTTs (0.3 g glucose per kg body weight with 0.02 U/kg insulin at 20 minutes; 30 samples over 180 minutes) performed in 40 healthy Caucasian subjects (male/female, 22/18; age, 46 +/- 9 years; body mass index [BMI], 26.7 +/- 5.7 kg. m(-2); mean +/- SD). The population methods assumed a log-normal population distribution of parameters. All methods gave a similar but overestimated population SG by 9% to 13%. Population SI was underestimated to a different degree by the methods (STS 6%, ITS 10%, MCMC 13%, and NM 7%). The between-subject variability of SG was overestimated by STS and underestimated by the population methods (true 33%, STS 40%, ITS 19%, MCMC 24%, NM 24%; coefficient of variation). For SI, this quantity was well estimated by all methods (true 79%, STS 80%, ITS 82%, MCMC 83%, NM 82%). The results for individual estimates indicate that STS performs better than the population methods when estimating SI (STS 12%, ITS 16%, MCMC 16%, NM 16%; 1 outlying subject excluded; root mean squared error expressed as percent of mean) but worse for SG (STS 28%, ITS 21%, MCMC 20%, NM 19%). We conclude that the robust performance of population approaches, preventing parameter estimation failures associated with the nonlinear regression analysis, is not required with IM-FSIVGTT in subjects with normal glucose tolerance. The standard two-stage technique is the preferred method under such circumstances.

Adult↗

Fine-scale mapping of disease loci via shattered coalescent modeling of genealogies.

We present a Bayesian, Markov-chain Monte Carlo method for fine-scale linkage-disequilibrium gene mapping using high-density marker maps. The method explicitly models the genealogy underlying a sample of case chromosomes in the vicinity of a putative disease locus, in contrast with the assumption of a star-shaped tree made by many existing multipoint methods. Within this modeling framework, we can allow for missing marker information and for uncertainty about the true underlying genealogy and the makeup of ancestral marker haplotypes. A crucial advantage of our method is the incorporation of the shattered coalescent model for genealogies, allowing for multiple founding mutations at the disease locus and for sporadic cases of disease. Output from the method includes approximate posterior distributions of the location of the disease locus and population-marker haplotype proportions. In addition, output from the algorithm is used to construct a cladogram to represent genetic heterogeneity at the disease locus, highlighting clusters of case chromosomes sharing the same mutation. We present detailed simulations to provide evidence of improvements over existing methodology. Furthermore, inferences about the location of the disease locus are shown to remain robust to modeling assumptions.

Algorithms↗

A method for evaluating the results of Bayesian model selection: application to linkage analyses of attributes determined by two or more genes.

OBJECTIVES: We apply and evaluate the intrinsic Bayes factor (IBF) of Berger and Pericchi [J Am Stat Assoc 1996;91:109-122; Bayesian Statistics, Oxford University Press, vol 5, 1996] to linkage analyses done using the stochastic search variable selection (SSVS) method of George and McCulloch [J Am Stat Assoc 1993;88:881-889] as proposed by Suh et al. [Genet Epidemiol 2001;21(suppl 1):S706-S711]. METHODS: We consider 20 simulations of linkage data obtained under two different generating models. The SSVS is applied to a multiple regression extension [Genet Epidemiol 2001;21(suppl 1): S706-S711] of the Haseman-Elston [Behav Genet 1972;2:3-19; Genet Epidemiol 2000;19:1-17] methods. Four prior distributions are considered. We apply the IBF criterion to those samples where different prior distributions result in different top models. RESULTS: In those samples where three different models were obtained using the four priors, application of the IBFs eliminated one of the two wrong models in 4 out of 5 situations. Further elimination using the IBF criterion for situations with two different subsets did not serve as well. CONCLUSIONS: When different priors result in three or more different subsets of markers, one can use the IBF to get this number down to two for consideration. When two subsets result we recommend that both be considered.

Bayes Theorem↗