Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Are any primroses (Primula) primitively monomorphic?

Primula (c. 430 species) and relatives (Primulaceae) are paradigmatic to our understanding of distyly. However, the common co-occurrence of distyly and monomorphy in closely related groups within the family has made the interpretation of its evolution difficult.Here, we infer a chloroplast DNA (cpDNA) phylogeny for 207 accessions, including 51% of the species and 95% of the sections of Primula with monomorphic populations, using Bayesian methods. With this tree, we infer the distribution of ancestral states on critical nodes using parsimony and likelihood methods. The inferred cpDNA phylogeny is consistent with prior estimates. The most recent common ancestor (MRCA) of Primula is resolved as distylous using both methods of inference. However, whether the distyly in Primula, Hottonia, and Vitaliana arose once or three independent times is not clear. We conclude that monomorphism in descendants of the MRCA of Primula is derived from distyly in all cases. Thus, scenarios for the evolution of distyly that rely on the persistence of primitive monomorphy (such as in Primula section Sphondylia) require re-evaluation.

DNA, Chloroplast↗

A phylogeny of the schoenoid sedges (Cyperaceae: Schoeneae) based on plastid DNA sequences, with special reference to the genera found in Africa.

Despite its large size (about 700 species), the australy-centred sedge tribe Schoeneae has received little explicit phylogenetic study, especially using molecular data. As a result, generic relationships are poorly understood, and even the monophyly of the tribe is open to question. In this study, plastid DNA sequences (rbcL, trnL-trnF, and rps16) drawn from a broad array of Schoeneae are analysed using Bayesian and parsimony-based approaches to infer a framework phylogeny for the tribe. Both analytical methods broadly support the monophyly of Schoeneae, Bayesian methods doing so with good support. Within the schoenoid clade, there is strong support for a series of monophyletic generic groupings whose interrelationships are unclear. These lineages form a large polytomy at the base of Schoeneae that may be indicative of past radiation, probably following the fragmentation of Gondwana. Most of these lineages contain both African and non-African members, suggesting a history of intercontinental dispersal. The results of this study clearly identify the relationships of the African-endemic schoenoid genera and demonstrate that the African-Australasian genus Tetraria, like Costularia, is polyphyletic. This pattern is morphologically consistent and suggests that these genera require realignment.

Africa↗

Applications of Bayesian statistical methods in microarray data analysis.

Microarray technology allows one to measure gene expression levels simultaneously on the whole-genome scale. The rapid progress generates both a great wealth of information and challenges in making inferences from such massive data sets. Bayesian statistical modeling offers an alternative approach to frequentist methodologies, and has several features that make these methods advantageous for the analysis of microarray data. These include the incorporation of prior information, flexible exploration of arbitrarily complex hypotheses, easy inclusion of nuisance parameters, and relatively well developed methods to handle missing data. Recent developments in Bayesian methodology generated a variety of techniques for the identification of differentially expressed genes, finding genes with similar expression profiles, and uncovering underlying gene regulatory networks. Bayesian methods will undoubtedly become more common in the future because of their great utility in microarray analysis.

Bayes Theorem↗

H-CORE: enabling genome-scale Bayesian analysis of biological systems without prior knowledge.

The Bayesian network is a popular tool for describing relationships between data entities by representing probabilistic (in)dependencies with a directed acyclic graph (DAG) structure. Relationships have been inferred between biological entities using the Bayesian network model with high-throughput data from biological systems in diverse fields. However, the scalability of those approaches is seriously restricted because of the huge search space for finding an optimal DAG structure in the process of Bayesian network learning. For this reason, most previous approaches limit the number of target entities or use additional knowledge to restrict the search space. In this paper, we use the hierarchical clustering and order restriction (H-CORE) method for the learning of large Bayesian networks by clustering entities and restricting edge directions between those clusters, with the aim of overcoming the scalability problem and thus making it possible to perform genome-scale Bayesian network analysis without additional biological knowledge. We use simulations to show that H-CORE is much faster than the widely used sparse candidate method, whilst being of comparable quality. We have also applied H-CORE to retrieving gene-to-gene relationships in a biological system (The 'Rosetta compendium'). By evaluating learned information through literature mining, we demonstrate that H-CORE enables the genome-scale Bayesian analysis of biological systems without any prior knowledge.

Algorithms↗

Model-independent mean-field theory as a local method for approximate propagation of information.

We present a systematic approach to mean-field theory (MFT) in a general probabilistic setting without assuming a particular model. The mean-field equations derived here may serve as a local, and thus very simple, method for approximate inference in probabilistic models such as Boltzmann machines or Bayesian networks. Our approach is 'model-independent' in the sense that we do not assume a particular type of dependences; in a Bayesian network, for example, we allow arbitrary tables to specify conditional dependences. In general, there are multiple solutions to the mean-field equations. We show that improved estimates can be obtained by forming a weighted mixture of the multiple mean-field solutions. Simple approximate expressions for the mixture weights are given. The general formalism derived so far is evaluated for the special case of Bayesian networks. The benefits of taking into account multiple solutions are demonstrated by using MFT for inference in a small and in a very large Bayesian network. The results are compared with the exact results.

Child↗

A Bayesian framework for the analysis of microarray expression data: regularized t -test and statistical inferences of gene changes.

MOTIVATION: DNA microarrays are now capable of providing genome-wide patterns of gene expression across many different conditions. The first level of analysis of these patterns requires determining whether observed differences in expression are significant or not. Current methods are unsatisfactory due to the lack of a systematic framework that can accommodate noise, variability, and low replication often typical of microarray data. RESULTS: We develop a Bayesian probabilistic framework for microarray data analysis. At the simplest level, we model log-expression values by independent normal distributions, parameterized by corresponding means and variances with hierarchical prior distributions. We derive point estimates for both parameters and hyperparameters, and regularized expressions for the variance of each gene by combining the empirical variance with a local background variance associated with neighboring genes. An additional hyperparameter, inversely related to the number of empirical observations, determines the strength of the background variance. Simulations show that these point estimates, combined with a t -test, provide a systematic inference approach that compares favorably with simple t -test or fold methods, and partly compensate for the lack of replication.

Bayes Theorem↗

A mixture model for bovine abortion and foetal survival.

The effect of spontaneous abortion on the dairy industry is substantial, costing the industry on the order of US dollars 200 million per year in California alone. We analyse data from a cohort study of nine dairy herds in Central California. A key feature of the analysis is the observation that only a relatively small proportion of cows will abort (around 10;15 per cent), so that it is inappropriate to analyse the time-to-abortion (TTA) data as if it were standard censored survival data, with cows that fail to abort by the end of the study treated as censored observations. We thus broaden the scope to consider the analysis of foetal lifetime distribution (FLD) data for the cows, with the dual goals of characterizing the effects of various risk factors on (i). the likelihood of abortion and, conditional on abortion status, on (ii). the risk of early versus late abortion. A single model is developed to accomplish both goals with two sets of specific herd effects modelled as random effects. Because multimodal foetal hazard functions are expected for the TTA data, both a parametric mixture model and a non-parametric model are developed. Furthermore, the two sets of analyses are linked because of anticipated dependence between the random herd effects. All modelling and inferences are accomplished using modern Bayesian methods.

Abortion, Veterinary↗

Examining accuracy of screening mammography using an event order model.

Screening mammography is a widely used method for breast cancer detection. For each mammogram we propose a performance model based on order of outcomes. That is, we envision an initial assessment, a follow up assessment if the initial one is positive and, eventually, a determination of whether cancer was present or not. A model can be built at each stage reflecting effects due to patient characteristics, to the facility where mammogram was performed and to the radiologist reading the mammogram. Since assessment is not perfectly associated with outcome, familiar rates of agreement and disagreement are of interest. These rates can be investigated at various levels of risk factors of interest. The approach is illustrated with screening mammography data from the Group Health Cooperative in Seattle, WA. A Bayesian framework is adopted for inference and an analysis of the data set is presented.

Adult↗

A cluster model for space-time disease counts.

Modelling disease clustering over space and time can be helpful in providing indications of possible exposures and planning corresponding public health practices. Though a considerable number of studies focus on modelling spatio-temporal patterns of disease, most of them do not directly model a spatio-temporal clustering structure and could be ineffective for detecting clusters. In this paper, we extend a purely spatial cluster model to accommodate space-time clustering. Inference is performed in a Bayesian framework using reversible jump Markov chain Monte Carlo. This idea is illustrated using data on female breast cancer mortality from Japan. A hierarchical parametric space-time model for mapping disease is used for comparison.

Adult↗

Estimation of the proportion of overweight individuals in small areas - a robust extension of the Fay-Herriot model.

Hierarchical model such as Fay-Herriot (FH) model is often used in small area estimation. The method might perform well overall but is vulnerable to outliers. We propose a robust extension of the FH model by assuming the area random effects follow a t distribution with an unknown degrees-of-freedom parameter. The inferences are constructed using a Bayesian framework. Monte Carlo Markov Chain (MCMC) such as Gibbs sampling and Metropolis-Hastings acceptance and rejection algorithms are used to obtain the joint posterior distribution of model parameters. The procedure is used to estimate the county-level proportion of overweight individuals from the 2003 public-use Behavioral Risk Factor Surveillance System (BRFSS) data. We also discuss two approaches for identifying outliers in the context of this application.

Adolescent↗

Mitochondrial phylogeny of Anura (Amphibia): a case study of congruent phylogenetic reconstruction using amino acid and nucleotide characters.

We explore whether phylogenetic analyses of the same sequence data set at the amino acid and nucleotide level are able to recover congruent topologies, as well as the advantages and limitations of both alternative approaches. As a case study, mitochondrial protein-coding genes were used to discern among competing hypotheses on the phylogenetic relationships of major anuran amphibian lineages. To properly address this phylogenetic question, the complete nucleotide sequences of the mitochondrial genomes of two archaeobatrachian species, Ascaphus truei and Pelobates cultripes, were determined anew. Bayesian and maximum likelihood phylogenetic inferences of the same sequence data set were performed based on both amino acid and nucleotide characters, with the latter analysed either as codons or as a reduced data set of first+second (P12) codon positions. In addition, likelihood-based ratio tests were performed to evaluate the support of alternative topologies. The different data sets arrived at congruent and highly supported topologies, suggesting a similar phylogenetic resolving power of the two character types provided that correctly selected sites and appropriate evolutionary models are used. The reconstructed anuran mitochondrial phylogeny supports the paraphyly of Archaeobatrachia, with Ascaphus as sister group to all the remaining anurans, and Pelobates as sister group of Neobatrachia. However, the employed tree reconstruction methods and likelihood-based ratio tests seemed to be negatively affected by the fast evolving sequences of neobatrachians, suggesting that the phylogeny of Anura here presented is not definitive, and needs further investigation using an extended taxon sampling.

Amphibian Proteins↗

Bayesian statistics for parasitologists.

Bayesian statistical methods are increasingly being used in the analysis of parasitological data. Here, the basis of differences between the Bayesian method and the classical or frequentist approach to statistical inference is explained. This is illustrated with practical implications of Bayesian analyses using prevalence estimation of strongyloidiasis and onchocerciasis as two relevant examples. The strongyloidiasis example addresses the problem of parasitological diagnosis in the absence of a gold standard, whereas the onchocerciasis case focuses on the identification of villages warranting priority mass ivermectin treatment. The advantages and challenges faced by users of the Bayesian approach are also discussed and the readers pointed to further directions for a more in-depth exploration of the issues raised. We advocate collaboration between parasitologists and Bayesian statisticians as a fruitful and rewarding venture for advancing applied research in parasite epidemiology and the control of parasitic infections.

Animals↗

Inferring the phylogenetic position of Boa constrictor among the Boinae.

Snakes of the subfamily Boinae are found in Madagascar, the Papuan-Pacific Islands, and the Neotropics. It has been suggested that genera within each of these particular areas do not form monophyletic groups. Further, it was proposed that the New World Boa constrictor is more closely related to boine genera in Madagascar than to boines in the Neotropics. Along with inferring the relationship of all boine genera using data from the cytochrome b gene and morphology, the placement of Boa was also examined. Phylogenetic inferences using maximum likelihood and Bayesian (BI) methods for combined data analyses and separate analyses of DNA sequence and morphological data were conducted. Priors, parametric bootstraps, and the Shimodaira-Hasegawa test were used to examine the previously proposed placement of Boa with Madagascan taxa using these DNA data. DNA data and combined data analyses strongly reject the hypothesis that Boa is more closely related to Old World genera than to other New World genera. Additionally, strong tree support suggests that all species within Madagascar, the Papuan-Pacific Islands, and the Neotropics each form a monophyletic group with respect to their geographic region.

Animals↗

A critique of Oaksford, Chater, and Larkin's (2000) conditional probability model of conditional reasoning.

M. Oaksford, N. Chater, and J. Larkin (2000) proffered a Bayesian model in which conditional inferences are a direct function of conditional probabilities. In the current article, the authors first considered this model regarding the processing of negatives in conditional reasoning. Its predictions were evaluated against a large-scale meta-analysis (W. J. Schroyens, W. Schaeken, & G. d'Ydewalle, 2001b). This evaluation shows that the model is flawed: The relative size of the negative effects does not match predictions. Next, the authors evaluated the model in relation to inferences about affirmative conditionals, again considering the results of a meta-analysis (W. J. Schroyens, W. Schaeken, & G. d'Ydewalle, 2001a). The conditional probability model is countered by the data reported in literature; a mental models based model produces a better fit. The authors conclude that a purely probabilistic model is deficient and incomplete and cannot do without algorithmic processing assumptions if it is to advance toward a descriptively adequate psychological theory.

Conditioning, Psychological↗

Detection of gene copy number changes in CGH microarrays using a spatially correlated mixture model.

MOTIVATION: Comparative genomic hybridization array experiments that investigate gene copy number changes present new challenges for statistical analysis and call for methods that incorporate spatial dependence between sequences along the chromosome. For this purpose, we propose a novel method called CGHmix. It is based on a spatially structured mixture model with three states corresponding to genomic sequences that are either unmodified, deleted or amplified. Inference is performed in a Bayesian framework. From the output, posterior probabilities of belonging to each of the three states are estimated for each genomic sequence and used to classify them. RESULTS: Using simulated data, CGHmix is validated and compared with both a conventional unstructured mixture model and with a recently proposed data mining method. We demonstrate the good performance of CGHmix for classifying copy number changes. In addition, the method provides a good estimate of the false discovery rate. We also present the analysis of a cancer related dataset. SUPPLEMENTARY INFORMATION: http://www.bgx.org.uk/papers.html

Algorithms↗

Bayesian statistical analyses for presence of single genes affecting meat quality traits in a crossed pig population.

Presence of single genes affecting meat quality traits was investigated in F2 individuals of a cross between Chinese Meishan and Western pig lines using phenotypic measurements on 11 traits. A Bayesian approach was used for inference about a mixed model of inheritance, postulating effects of polygenic background genes, action of a biallelic autosomal single gene and various nongenetic effects. Cooking loss, drip loss, two pH measurements, intramuscular fat, shearforce and back-fat thickness were traits found to be likely influenced by a single gene. In all cases, a recessive allele was found, which likely originates from the Meishan breed and is absent in the Western founder lines. By studying associations between genotypes assigned-to individuals based on phenotypic measurements for various traits, it was concluded that cooking loss, two pH measurements and possibly backfat thickness are influenced by one gene, and that a second gene influences intramuscular fat and possibly shearforce and drip loss. Statistical findings were supported by demonstrating marked differences in variances of families of fathers inferred as carriers and those inferred as noncarriers. It is concluded that further molecular genetic research effort to map single genes affecting these traits based on the same experimental data has a high probability of success.

Animals↗

Integrated surface model optimization for freehand three-dimensional echocardiography.

The major obstacle of three-dimensional (3-D) echocardiography is that the ultrasound image quality is too low to reliably detect features locally. Almost all available surface-finding algorithms depend on decent quality boundaries to get satisfactory surface models. We formulate the surface model optimization problem in a Bayesian framework, such that the inference made about a surface model is based on the integration of both the low-level image evidence and the high-level prior shape knowledge through a pixel class prediction mechanism. We model the probability of pixel classes instead of making explicit decisions about them. Therefore, we avoid the unreliable edge detection or image segmentation problem and the pixel correspondence problem. An optimal surface model best explains the observed images such that the posterior probability of the surface model for the observed images is maximized. The pixel feature vector as the image evidence includes several parameters such as the smoothed grayscale value and the minimal second directional derivative. Statistically, we describe the feature vector by the pixel appearance probability model obtained by a nonparametric optimal quantization technique. Qualitatively, we display the imaging plane intersections of the optimized surface models together with those of the ground-truth surfaces reconstructed from manual delineations. Quantitatively, we measure the projection distance error between the optimized and the ground-truth surfaces. In our experiment, we use 20 studies to obtain the probability models offline. The prior shape knowledge is represented by a catalog of 86 left ventricle surface models. In another set of 25 test studies, the average epicardial and endocardial surface projection distance errors are 3.2 +/- 0.85 mm and 2.6 +/- 0.78 mm, respectively.

Algorithms↗

Bayesian multivariate logistic regression.

Bayesian analyses of multivariate binary or categorical outcomes typically rely on probit or mixed effects logistic regression models that do not have a marginal logistic structure for the individual outcomes. In addition, difficulties arise when simple noninformative priors are chosen for the covariance parameters. Motivated by these problems, we propose a new type of multivariate logistic distribution that can be used to construct a likelihood for multivariate logistic regression analysis of binary and categorical data. The model for individual outcomes has a marginal logistic structure, simplifying interpretation. We follow a Bayesian approach to estimation and inference, developing an efficient data augmentation algorithm for posterior computation. The method is illustrated with application to a neurotoxicology study.

Algorithms↗