Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Model inference or model selection: discussion of Klugkist, Laudy, and Hoijtink (2005).

I. Klugkist, O. Laudy, and H. Hoijtink (2005) presented a Bayesian approach to analysis of variance models with inequality constraints. Constraints may play 2 distinct roles in data analysis. They may represent prior information that allows more precise inferences regarding parameter values, or they may describe a theory to be judged against the data. In the latter case, the authors emphasized the use of Bayes factors and posterior model probabilities to select the best theory. One difficulty is that interpretation of the posterior model probabilities depends on which other theories are included in the comparison. The posterior distribution of the parameters under an unconstrained model allows one to quantify the support provided by the data for inequality constraints without requiring the model selection framework.

Analysis of Variance↗

A Bayesian approach to modeling dynamic effective connectivity with fMRI data.

A state-space modeling approach for examining dynamic relationship between multiple brain regions was proposed in Ho, Ombao and Shumway (Ho, M.R., Ombao, H., Shumway, R., 2005. A State-Space Approach to Modelling Brain Dynamics to Appear in Statistica Sinica). Their approach assumed that the quantity representing the influence of one neuronal system over another, or effective connectivity, is time-invariant. However, more and more empirical evidence suggests that the connectivity between brain areas may be dynamic which calls for temporal modeling of effective connectivity. A Bayesian approach is proposed to solve this problem in this paper. Our approach first decomposes the observed time series into measurement error and the BOLD (blood oxygenation level-dependent) signals. To capture the complexities of the dynamic processes in the brain, region-specific activations are subsequently modeled, as a linear function of the BOLD signals history at other brain regions. The coefficients in these linear functions represent effective connectivity between the regions under consideration. They are further assumed to follow a random walk process so to characterize the dynamic nature of brain connectivity. We also consider the temporal dependence that may be present in the measurement errors. ML-II method (Berger, J.O., 1985. Statistical Decision Theory and Bayesian Analysis (2nd ed.). Springer, New York) was employed to estimate the hyperparameters in the model and Bayes factor was used to compare among competing models. Statistical inference of the effective connectivity coefficients was based on their posterior distributions and the corresponding Bayesian credible regions (Carlin, B.P., Louis, T.A., 2000. Bayes and Empirical Bayes Methods for Data Analysis (2nd ed.). Chapman and Hall, Boca Raton). The proposed method was applied to a functional magnetic resonance imaging data set and results support the theory of attentional control network and demonstrate that this network is dynamic in nature.

Attention↗

Comparison of Bayesian model averaging and stepwise methods for model selection in logistic regression.

Logistic regression is the standard method for assessing predictors of diseases. In logistic regression analyses, a stepwise strategy is often adopted to choose a subset of variables. Inference about the predictors is then made based on the chosen model constructed of only those variables retained in that model. This method subsequently ignores both the variables not selected by the procedure, and the uncertainty due to the variable selection procedure. This limitation may be addressed by adopting a Bayesian model averaging approach, which selects a number of all possible such models, and uses the posterior probabilities of these models to perform all inferences and predictions. This study compares the Bayesian model averaging approach with the stepwise procedures for selection of predictor variables in logistic regression using simulated data sets and the Framingham Heart Study data. The results show that in most cases Bayesian model averaging selects the correct model and out-performs stepwise approaches at predicting an event of interest.

Age Factors↗

Hierarchical bayesian models for regularisation in sequential learning

We show that a hierarchical Bayesian modeling approach allows us to perform regularization in sequential learning. We identify three inference levels within this hierarchy: model selection, parameter estimation, and noise estimation. In environments where data arrive sequentially, techniques such as cross validation to achieve regularization or model selection are not possible. The Bayesian approach, with extended Kalman filtering at the parameter estimation level, allows for regularization within a minimum variance framework. A multilayer perceptron is used to generate the extended Kalman filter nonlinear measurements mapping. We describe several algorithms at the noise estimation level that allow us to implement on-line regularization. We also show the theoretical links between adaptive noise estimation in extended Kalman filtering, multiple adaptive learning rates, and multiple smoothing regularization coefficients.

Journal Article↗

Flexible empirical Bayes models for differential gene expression.

MOTIVATION: Inference about differential expression is a typical objective when analyzing gene expression data. Recently, Bayesian hierarchical models have become increasingly popular for this type of problem. The two most common hierarchical models are the hierarchical Gamma-Gamma (GG) and Lognormal-Normal (LNN) models. However, to facilitate inference, some unrealistic assumptions have been made. One such assumption is that of a common coefficient of variation across genes, which can adversely affect the resulting inference. RESULTS: In this paper, we extend both the GG and LNN modeling frameworks to allow for gene-specific variances and propose EM based algorithms for parameter estimation. The proposed methodology is evaluated on three experimental datasets: one cDNA microarray experiment and two Affymetrix spike-in experiments. The two extended models significantly reduce the false positive rate while keeping a high sensitivity when compared to the originals. Finally, using a simulation study we show that the new frameworks are also more robust to model misspecification. AVAILABILITY: The R code for implementing the proposed methodology can be downloaded at http://www.stat.ubc.ca/~c.lo/FEBarrays. SUPPLEMENTARY INFORMATION: The supplementary material is available at http://www.stat.ubc.ca/~c.lo/FEBarrays/supp.pdf.

Algorithms↗

Bayesian estimation of positively selected sites.

In protein-coding DNA sequences, historical patterns of selection can be inferred from amino acid substitution patterns. High relative rates of nonsynonymous to synonymous changes (omega = dN/dS) are a clear indicator of positive, or directional, selection, and several recently developed methods attempt to distinguish these sites from those under neutral or purifying selection. One method uses an empirical Bayesian framework that accounts for varying selective pressures across sites while conditioning on the parameters of the model of DNA evolution and on the phylogenetic history. We describe a method that identifies sites under diversifying selection using a fully Bayesian framework. Similar to earlier work, the method presented here allows the rate of nonsynonymous to synonymous changes to vary among sites. The significant difference in using a fully Bayesian approach lies in our ability to account for uncertainty in parameters including the tree topology, branch lengths, and the codon model of DNA substitution. We demonstrate the utility of the fully Bayesian approach by applying our method to a data set of the vertebrate beta-globin gene. Compared to a previous analysis of this data set, the hierarchical model found most of the same sites to be in the positive selection class, but with a few striking exceptions.

Animals↗

Quality control in nerve conduction studies with coupled knowledge-based system approach.

Contemporary equipment used for nerve conduction studies is usually capable of computerized measurement of latency, amplitude, duration, and area of nerve and muscle action potentials and resulting conduction velocities. Abnormalities can be due to technical error or disease. Identification of technical error is a major element of quality control in electromyography, and artificial intelligence could be useful for this purpose. We have developed a coupled knowledge-based prototype system (QUALICON) to assess the correctness of recording and stimulating characteristics in routine conduction studies. QUALICON extracts numeric features from CMAPs or SNAPs, which are translated into symbolic form to drive a Bayesian network. The network uses high-level knowledge to infer the quality of stimulating and recording electrode placement as well as polarity and stimulus strength making recommendations as to the likely technical error when abnormal potentials are detected. A preliminary assessment shows that QUALICON performs as well as manual assessment performed by professionals.

Action Potentials↗

Bayesian analysis of non-homogeneous Markov chains: application to mental health data.

In this paper we present a formal treatment of non-homogeneous Markov chains by introducing a hierarchical Bayesian framework. Our work is motivated by the analysis of correlated categorical data which arise in assessment of psychiatric treatment programs. In our development, we introduce a Markovian structure to describe the non-homogeneity of transition patterns. In doing so, we introduce a logistic regression set-up for Markov chains and incorporate covariates in our model. We present a Bayesian model using Markov chain Monte Carlo methods and develop inference procedures to address issues encountered in the analyses of data from psychiatric treatment programs. Our model and inference procedures are implemented to some real data from a psychiatric treatment study.

Adolescent↗

Investigation of major gene for milk yield, milking speed, dry matter intake, and body weight in dairy cattle.

The main aim of this study was to determine if there exist any major gene for milk yield (MY), milking speed (MS), dry matter intake (DMI), and body weight (BW) recorded at various stages of lactation in first-lactation dairy cows (2543 observations from 320 cows) kept at the research farm of the Swiss Federal Institute of Technology between April 1994 and April 2004. Data were modelled based a simple repeatability covariance structure and analysed by using Bayesian segregation analyses. Gibbs sampling was used to make statistical inferences on posterior distributions; inferences were based on a single run of the Markov chain for each trait with 500,000 samples, with each 10th sample collected because of the high correlation among the samples. The posterior mean (+/-SD) of major gene variance was 2.61 (+/-2.46) for MY, 0.83 (+/-1.26) for MS, 4.37 (+/-2.34) for DMI, and 2056.43 (+/-665.67) for BW. Highest posterior density regions for 3 of the 4 traits did not include 0 (except MS), which supported the evidence for major gene. With additional tests for agreement with Mendelian transmission probabilities, we could only confirm the existence of a major gene for MY, but not for MS, DMI, and BW. Expected Mendelian transmission probabilities and their model fits were also compared.

Animals↗

Sensitivity analysis of misclassification: a graphical and a Bayesian approach.

PURPOSE: Misclassification can produce bias in measures of association. Sensitivity analyses have been suggested to explore the impact of such bias, but do not supply formally justified interval estimates. METHODS: To account for exposure misclassification, recently developed Bayesian approaches were extended to incorporate prior uncertainty and correlation of sensitivity and specificity. Under nondifferential misclassification, a contour plot is used to depict relations among the corrected odds ratio, sensitivity, and specificity. RESULTS: Methods are illustrated by application to a case-control study of cigarette smoking and invasive pneumococcal disease while varying the distributional assumptions about sensitivity and specificity. Results are compared with those of conventional methods, which do not account for misclassification, and a sensitivity analysis, which assumes fixed sensitivity and specificity. CONCLUSION: By using Bayesian methods, investigators can incorporate uncertainty about misclassification into probabilistic inferences.

Bayes Theorem↗

Sciurid phylogeny and the paraphyly of Holarctic ground squirrels (Spermophilus).

The squirrel family, Sciuridae, is one of the largest and most widely dispersed families of mammals. In spite of the wide distribution and conspicuousness of this group, phylogenetic relationships remain poorly understood. We used DNA sequence data from the mitochondrial cytochrome b gene of 114 species in 21 genera to infer phylogenetic relationships among sciurids based on maximum parsimony and Bayesian phylogenetic methods. Although we evaluated more complex alternative models of nucleotide substitution to reconstruct Bayesian phylogenies, none provided a better fit to the data than the GTR+G+I model. We used the reconstructed phylogenies to evaluate the current taxonomy of the Sciuridae. At essentially all levels of relationships, we found the phylogeny of squirrels to be in substantial conflict with the current taxonomy. At the highest level, the flying squirrels do not represent a basal divergence, and the current division of Sciuridae into two subfamilies is therefore not phylogenetically informative. At the tribal level, the Neotropical pygmy squirrel, Sciurillus, represents a basal divergence and is not closely related to the other members of the tribe Sciurini. At the genus level, the sciurine genus Sciurus is paraphyletic with respect to the dwarf squirrels (Microsciurus), and the Holarctic ground squirrels (Spermophilus) are paraphyletic with respect to antelope squirrels (Ammospermophilus), prairie dogs (Cynomys), and marmots (Marmota). Finally, several species of chipmunks and Holarctic ground squirrels do not appear monophyletic, indicating a need for reevaluation of alpha taxonomy.

Animals↗

The genus Anthracoidea (Basidiomycota, Ustilaginales): a molecular phylogenetic approach using LSU rDNA sequences.

The phylogenetic relationship of 52 specimens representing 30 species of Anthracoidea (Ustilaginales) was investigated by molecular analyses using sequence data from the large subunit (LSU) of nuclear rDNA. Phylogenetic trees were inferred with neighbour-joining (NJ), maximum parsimony (MP), and Bayesian Markov chain Monte Carlo (MCMC) methods. The results are discussed with respect to the species concept and the subdivision of the genus into subgenera and sections. Collections from different hosts and localities were compared. Our analyses can neither support nor significantly reject the hypothesis of the bipartition of the genus Anthracoidea. Thus, the representatives of the subgenus Proceres appeared in the NJ analysis as a moderately supported monophylum, whereas MCMC analysis revealed a polyphyletic topology for this group. Paraphyly of the subgenus Anthracoidea was supported by all methods used. Sections Echinosporae and Leiosporae were each represented by two species in our analyses which grouped together with high support. Section Anthracoidea should be restricted to a highly supported group with extremely irregular to angular teliospore shape. However, these three sections do not cover the whole diversity of the subgenus Anthracoidea. Molecular data largely supported the traditional circumscription of species, and species delimitations are discussed.

Bayes Theorem↗

Protein structure elucidation from NMR proton densities.

The NMR-generated foc proton density affords a template to which the molecule has to be fitted to derive the structure. Here we present a computational protocol that achieves this goal. H(N) atoms are readily recognizable from (1)H/(2)H exchange or (1)H/(15)N heteronuclear single quantum correlation (HSQC) experiments. The primary structure is threaded through the unassigned foc by leapfrogging along peptidyl amide H(N)s and the connected H(alpha)s. Via a Bayesian approach, the probabilities of the sequential connectivity hypotheses are inferred from likelihoods of H(N)/H(N), H(N)/H(alpha), and H(alpha)/H(alpha) interatomic distances as well as (1)H NMR chemical shifts, both derived from public databases. Once the polypeptide sequence is identified, directionality becomes established, and the foc N and C termini are recognized. After a similar procedure, side chain H atoms are found, including discriminated cis/trans proline loci. The folded structure then is derived via a direct molecular dynamics embedding into mirror image-related representations of the foc and selected according to a lowest energy criterion. The method was applied to foc densities calculated for two protein domains, col 2 and kringle 2. The obtained structures are within 1.0-1.5 A (backbone heavy atoms) and 1.5-2.0 A (all heavy atoms) rms deviations from reported x-ray and/or NMR structures.

Computer Simulation↗

Mixed model analysis of quantitative trait loci.

We develop a mixed model approach of quantitative trait locus (QTL) mapping for a hybrid population derived from the crosses of two or more distinguished outbred populations. Under the mixed model, we treat the mean allelic value of each source population as the fixed effect and the allelic deviations from the mean as random effects so that we can partition the total genetic variance into between- and within-population variances. Statistical inference of the QTL parameters is obtained by using the Bayesian method implemented by Markov chain Monte Carlo (MCMC). This unified QTL mapping algorithm treats the fixed and random model approaches as special cases of the general mixed model methodology. Utility and flexibility of the method are demonstrated by using a set of simulated data.

Algorithms↗

DIGIT: a novel gene finding program by combining gene-finders.

We have developed a general purpose algorithm which finds genes by combining plural existing gene-finders. The algorithm has been implemented into a novel gene-finder named DIGIT. An outline of the algorithm is as follows. First, existing gene-finders are applied to an uncharacterized genomic sequence (input sequence). Next, DIGIT produces all possible exons from the results of gene-finders, and assigns them their exon types, reading frames and exon scores. Finally, DIGIT searches a set of exons whose additive score is maximized under their reading frame constraints. Bayesian procedure and a hidden Markov model are used to infer exon scores and search the exon set, respectively. We have designed DIGIT so as to combine the results of FGENESH, GENSCAN and HMMgene, and have assessed its prediction accuracy by using recently compiled benchmark data sets. For all data sets, DIGIT successfully discarded many false-positive exons predicted by individual gene-finders and yielded remarkable improvements in sensitivity and specificity at the gene level compared with the best gene level accuracies achieved by any single gene-finder.

Algorithms↗

Genetic change for clinical mastitis in Norwegian cattle: a threshold model analysis.

Records of clinical mastitis on 1.6 million first-lactation daughters of 2,411 Norwegian Cattle sires that were progeny tested from 1978 through 1998 were analyzed with a threshold model. The main objective was to infer genetic change for the disease in the population. A Bayesian approach via Gibbs sampling was used. The model for the underlying liability had age at first calving, month x year of calving, herd x 3-year-period, and sire of the cow as explanatory variables. Posterior mean (SD) of heritability of liability to clinical mastitis was 0.066 (0.003). Genetic evaluations (posterior means) of sires both in the liability and observable scales were computed. Annual genetic change of liability to clinical mastitis for progeny tested bulls born from 1973 to 1993 was assessed. The linear regression of mean sire effect on year of birth had a posterior mean (SD) of -0.00018 (0.0004), suggesting a nearly constant genetic level for clinical mastitis. However, an analysis of sire posterior means by birth-year of daughters indicated an approximately constant genetic level in the cow population from 1976 to 1990 (-0.02%/yr), and a genetic improvement thereafter (-0.27%/yr). This reflects more emphasis on mastitis in selection of bulls in recent years. Corresponding results obtained with a standard linear model analysis were -0.01% and -0.23% per year, respectively (regression of sire predicted transmitting ability on birth-year of daughters). Genetic change seems to be slightly understated with the linear model, assuming the threshold model holds true.

Animals↗

Practical implications of modes of statistical inference for causal effects and the critical role of the assignment mechanism.

Causal inference in an important topic and one that is now attracting serious attention of statisticians. Although there exist recent discussions concerning the general definition of causal effects and a substantial literature on specific techniques for the analysis of data in randomized and nonrandomized studies, there has been relatively little discussion of modes of statistical inference for causal effects. This presentation briefly describes and contrasts four basic modes of statistical inference for causal effects, emphasizes the common underlying causal framework with a posited assignment mechanism, and describes practical implications in the context of an example involving the effects of switching from a name-band to a generic drug. A fundamental conclusion is that in such nonrandomized studies, sensitivity of inference to the assignment mechanism is the dominant issue, and it cannot be avoided by changing modes of inference, for instance, by changing from randomization-based to Bayesian methods.

Bayes Theorem↗

The use of multiple restriction fragment length polymorphisms in prenatal risk estimation. I. X-linked diseases.

An analytical procedure for estimating the risk of X-linked diseases based on presence/absence of a series of restriction sites is presented. Multiple-locus linkage phase of the carrier mother is first inferred from previous offspring, from parents, and by molecular means. Bayesian risk estimates are then obtained using this information and the recombination-segregation distribution. The improvement afforded by using multiple flanking markers rather than a single marker is dramatic. Whereas the upper bound on the probability that a family will be informative using a single diallelic X-linked marker is .5, in the case of m markers, the bound on the probability of an informative family becomes 1 - .5m. With a single linked marker, the precision in the risk estimate is bounded by the frequency of recombination, whereas the requirement of very tight linkage is relaxed somewhat when multiple flanking markers are used. Recombination interference and multiple-locus linkage disequilibria can further improve the risk estimates, but it is important to understand how the statistical confidence in these parameters affects the reliability of the risk estimates.

Chromosome Mapping↗