Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Estimation of population parameters and recombination rates from single nucleotide polymorphisms.

Some general likelihood and Bayesian methods for analyzing single nucleotide polymorphisms (SNPs) are presented. First, an efficient method for estimating demographic parameters from SNPs in linkage equilibrium is derived. The method is applied in the estimation of growth rates of a human population based on 37 SNP loci. It is demonstrated how ascertainment biases, due to biased sampling of loci, can be avoided, at least in some cases, by appropriate conditioning when calculating the likelihood function. Second, a Markov chain Monte Carlo (MCMC) method for analyzing linked SNPs is developed. This method can be used for Bayesian and likelihood inference on linked SNPs. The utility of the method is illustrated by estimating recombination rates in a human data set containing 17 SNPs and 60 individuals. Both methods are based on assumptions of low mutation rates.

Genetic Linkage↗

False-positive selection identified by ML-based methods: examples from the Sig1 gene of the diatom Thalassiosira weissflogii and the tax gene of a human T-cell lymphotropic virus.

Sexually induced gene 1 (Sig1) in the centric diatom Thalassiosira weissflogii is considered to encode a gamete recognition protein. Sorhannus (2003) analyzed nucleotide sequences of Sig1 using parsimony analysis and the maximum-likelihood (ML)-based Bayesian method for inferring positive selection at single amino acid sites and reported that positively selected sites were detected by the latter method but not by the former. He then concluded that for this type of study, the ML-based method is more reliable than parsimony analysis. Here we show that his results apparently represent false-positive cases of the ML-based method and that there is no solid evidence that this gene contains positively selected sites. We further demonstrate that in the tax gene of human T-cell lymphotropic virus type I (HTLV-I), all codon sites, including invariable sites, can be inferred as positively selected sites by the ML-based method. These observations indicate that the ML-based method may produce many false-positive sites. One of the main reasons for the occurrence of false positives is that in the ML-based method, codon sites are grouped into several categories, with different nonsynonymous/synonymous rate ratios (omegas), on a purely statistical basis, and positive selection is inferred indirectly by examining whether the average omega for each category is greater than 1. In parsimony analysis, however, the evolutionary change of nucleotides at each codon site is examined. For this reason, parsimony-based methods rarely produce false positives and are safer than ML-based methods for detecting positive selection at individual codon sites, although a large number of sequences are necessary.

Bayes Theorem↗

Bayesian estimation of disease prevalence and the parameters of diagnostic tests in the absence of a gold standard.

It is common in population screening surveys or in the investigation of new diagnostic tests to have results from one or more tests investigating the same condition or disease, none of which can be considered a gold standard. For example, two methods often used in population-based surveys for estimating the prevalence of a parasitic or other infection are stool examinations and serologic testing. However, it is known that results from stool examinations generally underestimate the prevalence, while serology generally results in overestimation. Using a Bayesian approach, simultaneous inferences about the population prevalence and the sensitivity, specificity, and positive and negative predictive values of each diagnostic test are possible. The methods presented here can be applied to each test separately or to two or more tests combined. Marginal posterior densities of all parameters are estimated using the Gibbs sampler. The techniques are applied to the estimation of the prevalence of Strongyloides infection and to the investigation of the diagnostic test properties of stool examinations and serologic testing, using data from a survey of all Cambodian refugees who arrived in Montreal, Canada, during an 8-month period.

Bayes Theorem↗

Color vision of ancestral organisms of higher primates.

The color vision of mammals is controlled by photosensitive proteins called opsins. Most mammals have dichromatic color vision, but hominoids and Old World (OW) monkeys enjoy trichromatic vision, having the blue-, green-, and red-sensitive opsin genes. Most New World (NW) monkeys are either dichromatic or trichromatic, depending on the sex and genotype. Trichromacy in higher primates is believed to have evolved to facilitate the detection of yellow and red fruits against dappled foliage, but the process of evolutionary change from dichromacy to trichromacy is not well understood. Using the parsimony and the newly developed Bayesian methods, we inferred the amino acid sequences of opsins of ancestral organisms of higher primates. The results suggest that the ancestors of OW and NW monkeys lacked the green gene and that the green gene later evolved from the red gene. The fact that the red/green opsin gene has survived the long nocturnal stage of mammalian evolution and that it is under strong purifying selection in organisms that live in dark environments suggests that this gene has another important function in addition to color vision, probably the control of circadian rhythms.

Amino Acid Sequence↗

Population structure in African Drosophila melanogaster revealed by microsatellite analysis.

Tropical sub-Saharan regions are considered to be the geographical origin of Drosophila melanogaster. Starting from there, the species colonized the rest of the world after the last glaciation about 10 000 years ago. Consistent with this demographic scenario, African populations have been shown to harbour higher levels of microsatellite and sequence variation than cosmopolitan populations. Nevertheless, limited information is available on the genetic structure of African populations. We used X chromosomal microsatellite variation to study the population structure of D. melanogaster populations using 13 sampling sites in North, West and East Africa. These populations were compared to six European and one North American population. Significant population structure was found among African D. melanogaster populations. Using a Bayesian method for inferring population structure we detected two distinct groups of populations among African D. melanogaster. Interestingly, the comparison to cosmopolitan D. melanogaster populations indicated that one of the divergent African groups is closely related to cosmopolitan flies. Low, but significant levels of differentiation were observed for sub-Saharan D. melanogaster populations from West and East Africa.

Africa↗

Structured additive regression for categorical space-time data: a mixed model approach.

Motivated by a space-time study on forest health with damage state of trees as the response, we propose a general class of structured additive regression models for categorical responses, allowing for a flexible semiparametric predictor. Nonlinear effects of continuous covariates, time trends, and interactions between continuous covariates are modeled by penalized splines. Spatial effects can be estimated based on Markov random fields, Gaussian random fields, or two-dimensional penalized splines. We present our approach from a Bayesian perspective, with inference based on a categorical linear mixed model representation. The resulting empirical Bayes method is closely related to penalized likelihood estimation in a frequentist setting. Variance components, corresponding to inverse smoothing parameters, are estimated using (approximate) restricted maximum likelihood. In simulation studies we investigate the performance of different choices for the spatial effect, compare the empirical Bayes approach to competing methodology, and study the bias of mixed model estimates. As an application we analyze data from the forest health survey.

Bayes Theorem↗

Clinical significance not statistical significance: a simple Bayesian alternative to p values.

OBJECTIVES: To take the common "Bayesian" interpretation of conventional confidence intervals to its logical conclusion, and hence to derive a simple, intuitive way to interpret the results of public health and clinical studies. DESIGN AND SETTING: The theoretical basis and practicalities of the approach advocated is at first explained and then its use is illustrated by referring to the interpretation of a real historical cohort study. The study considered compared survival on haemodialysis (HD) with that on continuous ambulatory peritoneal dialysis (CAPD) in 389 patients dialysed for end stage renal disease in Leicestershire between 1974 and 1985. Careful interpretation of the study was essential. This was because although it had relatively low statistical power, it represented all of the data that were available at the time and it had to inform a critical clinical policy decision: whether or not to continue putting the majority of new patients onto CAPD. MEASUREMENTS AND ANALYSIS: Conventional confidence intervals are often interpreted using subjective probability. For example, 95% confidence intervals are commonly understood to represent a range of values within which one may be 95% certain that the true value of whatever one is estimating really lies. Such an interpretation is fundamentally incorrect within the framework of conventional, frequency-based, statistics. However, it is valid as a statement of Bayesian posterior probability, provided that the prior distribution that represents pre-existing beliefs is uniform, which means flat, on the scale of the main outcome variable. This means that there is a limited equivalence between conventional and Bayesian statistics, which can be used to draw simple Bayesian style statistical inferences from a standard analysis. The advantage of such an approach is that it permits intuitive inferential statements to be made that cannot be made within a conventional framework and this can help to ensure that logical decisions are taken on the basis of study results. In the particular practical example described, this approach is applied in the context of an analysis based upon proportional hazards (Cox) regression. MAIN RESULTS AND CONCLUSIONS: The approach proposed expresses conclusions in a manner that is believed to be a helpful adjunct to more conventional inferential statements. It is of greatest value in those situations in which statistical significance may bear little relation to clinical significance and a conventional analysis using p values is liable to be misleading. Perhaps most importantly, this includes circumstances in which an important public health or clinical decision must be based upon a study that has unavoidably low statistical power. However, it is also useful in situations in which a decision must be based upon a large study that indicates that an effect that is highly statistically significant seems too small to be of practical relevance. In the illustrative example described, the approach helped in making a decision regarding the use of CAPD in Leicestershire during the latter half of the 1980s.

Bayes Theorem↗

Online updating of space-time disease surveillance models via particle filters.

Online surveillance of disease has become an important issue in public health. In particular, the space-time monitoring of disease plays an important part in any syndromic system. However, methodology for these systems is generally lacking. One approach to space-time monitoring of health data is to consider the space-time model parameters as the focus and to monitor their changes as multivariate time series (Lawson AB. Some considerations in spatial-temporal analysis of public health surveillance data. In Brookmeyer R, Stroup DF eds. Monitoring the Health of Populations. Oxford University Press, 2004; Vidal Rodeiro CL, Lawson AB. Monitoring changes in spatio-temporal maps of disease. Biometrical Journal 2006; to appear). However with complex space-time models, this becomes very time consuming. Some simplifications may be necessary and these can be made in a number of ways. In this article, the focus is on particle filters that can be used to resample the history of the process and thereby reduce computation time. This article describes a particular case of particle filters, the resample-move algorithm, proposed by Gilks and Berzuini (Gilks WR, Berzuini C. Following a moving target--Monte Carlo inference for dynamic Bayesian models. Journal of the Royal Statistical Society, Series B 2001; 63: 127-46), in the context of disease map surveillance. This is followed by an application to a real data set in which a comparison between the use of Markov chain Monte Carlo methods and the resample-move algorithm is carried out.

Algorithms↗

Statistical analysis of nonlinear structural equation models with continuous and polytomous data.

A general nonlinear structural equation model with mixed continuous and polytomous variables is analysed. A Bayesian approach is proposed to estimate simultaneously the thresholds, the structural parameters and the latent variables. To solve the computational difficulties involved in the posterior analysis, a hybrid Markov chain Monte Carlo method that combines the Gibbs sampler and the Metropolis-Hasting algorithm is implemented to produce the Bayesian solution. Statistical inferences, which involve estimation of parameters and their standard errors, residuals and outliers analyses, and goodness-of-fit statistics for testing the posited model, are discussed. The proposed procedure is illustrated by a simulation study and a real example.

Humans↗

A fluctuation method to quantify in vivo fluorescence data.

Quantitative in vivo measurements are essential for developing a predictive understanding of cellular behavior. Here we present a technique that converts observed fluorescence intensities into numbers of molecules. By transiently expressing a fluorescently tagged protein and then following its dilution during growth and division, we observe asymmetric partitioning of fluorescence between daughter cells at each division. Such partition asymmetries are set by the actual numbers of proteins present, and thus provide a means to quantify fluorescence levels. We present a Bayesian algorithm that infers from such data both the fluorescence conversion factor and an estimate of the measurement error. Our algorithm works for arbitrarily sized data sets and handles consistently any missing measurements. We verify the algorithm with extensive simulation and demonstrate its application to experimental data from Escherichia coli. Our technique should provide a quantitative internal calibration to systems biology studies of both synthetic and endogenous cellular networks.

Algorithms↗

Phylogenetic mapping of recombination hotspots in human immunodeficiency virus via spatially smoothed change-point processes.

We present a Bayesian framework for inferring spatial preferences of recombination from multiple putative recombinant nucleotide sequences. Phylogenetic recombination detection has been an active area of research for the last 15 years. However, only recently attempts to summarize information from several instances of recombination have been made. We propose a hierarchical model that allows for simultaneous inference of recombination breakpoint locations and spatial variation in recombination frequency. The dual multiple change-point model for phylogenetic recombination detection resides at the lowest level of our hierarchy under the umbrella of a common prior on breakpoint locations. The hierarchical prior allows for information about spatial preferences of recombination to be shared among individual data sets. To overcome the sparseness of breakpoint data, dictated by the modest number of available recombinant sequences, we a priori impose a biologically relevant correlation structure on recombination location log odds via a Gaussian Markov random field hyperprior. To examine the capabilities of our model to recover spatial variation in recombination frequency, we simulate recombination from a predefined distribution of breakpoint locations. We then proceed with the analysis of 42 human immunodeficiency virus (HIV) intersubtype gag recombinants and identify a putative recombination hotspot.

Animals↗

Sequential updating of a new dynamic pharmacokinetic model for caffeine in premature neonates.

BACKGROUND AND OBJECTIVE: Caffeine treatment is widely used in nursing care to reduce the risk of apnoea in premature neonates. To check the therapeutic efficacy of the treatment against apnoea, caffeine concentration in blood is an important indicator. The present study was aimed at building a pharmacokinetic model as a basis for a medical decision support tool. METHODS: In the proposed model, time dependence of physiological parameters is introduced to describe rapid growth of neonates. To take into account the large variability in the population, the pharmacokinetic model is embedded in a population structure. The whole model is inferred within a Bayesian framework. To update caffeine concentration predictions as data of an incoming patient are collected, we propose a fast method that can be used in a medical context. This involves the sequential updating of model parameters (at individual and population levels) via a stochastic particle algorithm. RESULTS: Our model provides better predictions than the ones obtained with models previously published. We show, through an example, that sequential updating improves predictions of caffeine concentration in blood (reduce bias and length of credibility intervals). The update of the pharmacokinetic model using body mass and caffeine concentration data is studied. It shows how informative caffeine concentration data are in contrast to body mass data. CONCLUSION: This study provides the methodological basis to predict caffeine concentration in blood, after a given treatment if data are collected on the treated neonate.

Bayes Theorem↗

Ascoma morphology is homoplaseous and phylogenetically misleading in some pyrenocarpous lichens.

The phylogenetic relationships of many lichen-forming perithecioid ascomycetes are unknown. We generated nuLSU and mtSSU rDNA sequences of members of seven families of pyrenocarpous lichens and used a Bayesian framework to infer a phylogenetic estimate. Members of the perithecioid Protothelenellaceae, Thelenellaceae and Thrombiaceae surprisingly cluster within the mainly discocarpous Lecanoromycetes, while Strigulaceae, Verrucariaceae and Pyrenulaceae are related to the ascolocular Chaetothyriomycetes. Micromorphological studies of the ascomata showed that the two main groups of pyrenocarpous lichen-forming fungi differ in their ascus types. The Strigulaceae, Verrucariaceae and Pyrenulaceae have apically and laterally thick-walled asci, whereas the Thelenellaceae, Protothelenellaceae and Thrombiaceae have only apically thickened asci. The latter two show ring-shaped amyloid apical structures. Based on morphological and molecular evidence we propose to reduce Thrombiaceae to synonymy with Protothelenellaceae.

Ascomycota↗

Automated linkage of free-text descriptions of patients with a practice guideline.

The process of applying a practice guideline to a patient requires a great deal of clinical data. AAPT (Appropriateness-Assessment Processing from Text) is an experimental computer program that can assess the appropriateness of coronary-artery bypass grafting surgery (CABG) in patients with coronary-artery disease (CAD) and chronic stable angina from the admission summaries of those patients. The AAPT architecture combines natural-language processing (NLP) and probabilistic inference. The NLP module identifies single clinical concepts of interest in the free-text document. The probabilistic inference module, a Bayesian belief network, estimates values for variables not specifically mentioned. AAPT produces a patient's summary of CAD that is similar to a manually generated clinical summary. Work is ongoing to improve AAPT and evaluate it as a tool to assist in the dissemination of guidelines and as a tool to encourage adherence to practice guidelines.

Angina Pectoris↗

Analysis of a Bayesian repeated measures model for detecting differences in GP prescribing habits.

A linear mixed model is used to detect a change, if any, in the prescribing habits in the UK at the general practice (family medicine) level due to an educational intervention given repeated measures data before and after the intervention and a control group. Inferences are corrected for general practice size and fundholding status. The estimates of the model parameters are obtained using Bayesian inference by applying Gibbs sampling. We develop three different priors for the parameters of the model. These three priors correspond to 'sceptical,' 'reference' and 'enthusiastic' priors in terms of the opinion about the treatment effects that they represent. We compare the results obtained by using these three priors for the parameters in the random effects model.

Anti-Inflammatory Agents, Non-Steroidal↗

A phylogeny of the Australian Sphenomorphus group (Scincidae: Squamata) and the phylogenetic placement of the crocodile skinks (Tribolonotus): Bayesian approaches to assessing congruence and obtaining confidence in maximum likelihood inferred relationships.

Australian scincid lizards are a diverse squamate assemblage ( approximately 385 species), divided among three major clades (Egernia, Eugongylus, and Sphenomorphus groups). The Sphenomorphus group is the largest, comprising 61% of the Australian scincid fauna. Phylogenetic relationships within the Australian Sphenomorphus group and the phylogenetic placement of Tribolonotus are inferred using mtDNA (12S and 16S rRNA genes, ND4 protein-coding gene, and associated tRNA genes; 2185bp total). These data were analyzed separately (structural RNA vs protein-coding partitions) and combined using maximum likelihood. Confidence in inferred clades was assessed using non-parametric bootstrapping and Bayesian analysis. Analysis of the combined data strongly supports Sphenomorphus group (as well as the Australian subgroup) monophyly. Notoscincus is strongly placed as the sister taxon of the remaining Australian Sphenomorphus group taxa, with this more exclusive clade being divided into two major groups (one restricted to mesic eastern Australia and the other continent wide). The speciose Australian "Eulamprus" and "Glaphyromorphus" are both polyphyletic. All remaining non-Sphenomorphus group lygosomine skinks strongly form a clade, with Tribolonotus placed as the sister taxon of the Australian Egernia group.

Animals↗

Bayesian phylogenetics using an RNA substitution model applied to early mammalian evolution.

We study the phylogeny of the placental mammals using molecular data from all mitochondrial tRNAs and rRNAs of 54 species. We use probabilistic substitution models specific to evolution in base paired regions of RNA. A number of these models have been implemented in a new phylogenetic inference software package for carrying out maximum likelihood and Bayesian phylogenetic inferences. We describe our Bayesian phylogenetic method which uses a Markov chain Monte Carlo algorithm to provide samples from the posterior distribution of tree topologies. Our results show support for four primary mammalian clades, in agreement with recent studies of much larger data sets mainly comprising nuclear DNA. We discuss some issues arising when using Bayesian techniques on RNA sequence data.

Animals↗

Modularized learning of genetic interaction networks from biological annotations and mRNA expression data.

MOTIVATION: Inferring the genetic interaction mechanism using Bayesian networks has recently drawn increasing attention due to its well-established theoretical foundation and statistical robustness. However, the relative insufficiency of experiments with respect to the number of genes leads to many false positive inferences. RESULTS: We propose a novel method to infer genetic networks by alleviating the shortage of available mRNA expression data with prior knowledge. We call the proposed method 'modularized network learning' (MONET). Firstly, the proposed method divides a whole gene set to overlapped modules considering biological annotations and expression data together. Secondly, it infers a Bayesian network for each module, and integrates the learned subnetworks to a global network. An algorithm that measures a similarity between genes based on hierarchy, specificity and multiplicity of biological annotations is presented. The proposed method draws a global picture of inter-module relationships as well as a detailed look of intra-module interactions. We applied the proposed method to analyze Saccharomyces cerevisiae stress data, and found several hypotheses to suggest putative functions of unclassified genes. We also compared the proposed method with a whole-set-based approach and two expression-based clustering approaches.

Algorithms↗