Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

Assessing and comparing costs: how robust are the bootstrap and methods based on asymptotic normality?

This article addresses and challenges some common perceptions in the statistical assessment of costs and cost-effectiveness in health economics. Cost data typically exhibit highly skew distributions. Two techniques whose validity does not depend on any specific form of underlying distribution are the bootstrap and methods based on asymptotic normality of sample means. These methods are generally thought to be appropriate for the analysis of cost data. We argue that, even when these methods are technically valid, they may often lead to inefficient and even misleading inferences. It is important to apply methods that recognise the skewness in cost data. We further demonstrate that it may also be important to incorporate relevant prior information in a Bayesian analysis.

Bayes Theorem↗

Comparison of recent methods for inference of variable influence in neural networks.

Neural networks (NNs) belong to 'black box' models and therefore 'suffer' from interpretation difficulties. Four recent methods inferring variable influence in NNs are compared in this paper. The methods assist the interpretation task during different phases of the modeling procedure. They belong to information theory (ITSS), the Bayesian framework (ARD), the analysis of the network's weights (GIM), and the sequential omission of the variables (SZW). The comparison is based upon artificial and real data sets of differing size, complexity and noise level. The influence of the neural network's size has also been considered. The results provide useful information about the agreement between the methods under different conditions. Generally, SZW and GIM differ from ARD regarding the variable influence, although applied to NNs with similar modeling accuracy, even when larger data sets sizes are used. ITSS produces similar results to SZW and GIM, although suffering more from the 'curse of dimensionality'.

Algorithms↗

Learning yeast gene functions from heterogeneous sources of data using hybrid weighted Bayesian networks.

We developed a machine learning system for determining gene functions from heterogeneous sources of data sets using a Weighted Naive Bayesian Network (WNB). The knowledge of gene functions is crucial for understanding many fundamental biological mechanisms such as regulatory pathways, cell cycles and diseases. Our major goal is to accurately infer functions of putative genes or ORFs (Open Reading Frames) from existing databases using computational methods. However, this task is intrinsically difficult since the underlying biological processes represent complex interactions of multiple entities. Therefore many functional links would be missing when only one or two source of data is used in the prediction. Our hypothesis is that integrating evidence from multiple and complementary sources could significantly improve the prediction accuracy. In this paper, our experimental results not only suggest that the above hypothesis is valid, but also provide guidelines for using the WNB system for data collection, training and predictions. The combined training data sets contain information from gene annotations, gene expressions, clustering outputs, keyword annotations and sequence homology from public databases. The current system is trained and tested on the genes of budding yeast Saccharomyces cerevisiae. Our WNB model can also be used to analyze the contribution of each source of information toward the prediction performance through the weight training process. The contribution analysis could potentially lead to significant scientific discovery by facilitating the interpretation and understanding of the complex relationships between biological entities.

Artificial Intelligence↗

Joint learning of gene functions--a Bayesian network model approach.

In this paper, we develop a machine learning system for determining gene functions from heterogeneous data sources using a Weighted Naive Bayesian network (WNB). The knowledge of gene functions is crucial for understanding many fundamental biological mechanisms such as regulatory pathways, cell cycles and diseases. Our major goal is to accurately infer functions of putative genes or Open Reading Frames (ORFs) from existing databases using computational methods. However, this task is intrinsically difficult since the underlying biological processes represent complex interactions of multiple entities. Therefore, many functional links would be missing when only one or two sources of data are used in the prediction. Our hypothesis is that integrating evidence from multiple and complementary sources could significantly improve the prediction accuracy. In this paper, our experimental results not only suggest that the above hypothesis is valid, but also provide guidelines for using the WNB system for data collection, training and predictions. The combined training data sets contain information from gene annotations, gene expressions, clustering outputs, keyword annotations, and sequence homology from public databases. The current system is trained and tested on the genes of budding yeast Saccharomyces cerevisiae. Our WNB model can also be used to analyze the contribution of each source of information toward the prediction performance through the weight training process. The contribution analysis could potentially lead to significant scientific discovery by facilitating the interpretation and understanding of the complex relationships between biological entities.

Artificial Intelligence↗

Methodological aspects in the assessment of treatment effects in observational health outcomes studies.

Prospective observational studies, which provide information on the effectiveness of interventions in natural settings, may complement results from randomised clinical trials in the evaluation of health technologies. However, observational studies are subject to a number of potential methodological weaknesses, mainly selection and observer bias. This paper reviews and applies various methods to control for selection bias in the estimation of treatment effects and proposes novel ways to assess the presence of observer bias. We also address the issues of estimation and inference in a multilevel setting. We describe and compare the use of regression methods, propensity score matching, fixed-effects models incorporating investigator characteristics, and a multilevel, hierarchical model using Bayesian estimation techniques in the control of selection bias. We also propose to assess the existence of observer bias in observational studies by comparing patient- and investigator-reported outcomes. To illustrate these methods, we have used data from the SOHO (Schizophrenia Outpatient Health Outcomes) study, a large, prospective, observational study of health outcomes associated with the treatment of schizophrenia. The methods used to adjust for differences between treatment groups that could cause selection bias yielded comparable results, reinforcing the validity of the findings. Also, the assessment of observer bias did not show that it existed in the SOHO study. Observational studies, when properly conducted and when using adequate statistical methods, can provide valid information on the evaluation of health technologies.

Bayes Theorem↗

Phylogeny of the festucoid grasses of subtribe Loliinae and allies (Poeae, Pooideae) inferred from ITS and trnL-F sequences.

Analyses of ribosomal ITS and chloroplast trnL-F sequences provide phylogenetic reconstruction for the festucoids (Poeae: Loliinae), a group of temperate grasses with morphological and molecular affinities to the large genus Festuca. Parsimony and Bayesian analyses of the combined ITS/trnL-F dataset show Loliinae to be monophyletic but unresolved for a weakly supported clade of 'broad-leaved Festuca,' a well-supported clade of 'fine-leaved Festuca,' and Castellia. The first group includes subgenera Schenodorus, Drymanthele, Leucopoa, and Subulatae, and sections Subbulbosae, Scariosa, and Pseudoscariosa of Festuca, plus Lolium and Micropyropsis. The second group includes sections Festuca, Aulaxyper, Eskia, and Amphigenes of Festuca, plus Vulpia, Ctenopsis, Psilurus, Wangenheimia, Cutandia, Narduroides, and Micropyrum. Subtribes Dactylidinae and Cynosurinae/Parapholiinae are sister clades and are the closest relatives of Loliinae. Vulpia is polyphyletic within the 'fine-leaved' fescues as revealed by the two genome analyses. Lolium is resolved as monophyletic in the ITS and combined analyses, but unresolved in the trnL-F based tree. Conflict between the ITS and the trnL-F trees in the placement of several taxa suggests the possibility of past reticulation events, although lineage sorting and possible ITS paralogy cannot be ruled out.

Bayes Theorem↗

Why we still use our heads instead of formulas: toward an integrative approach.

This review begins with a discussion of Meehl's (1957) query regarding when to use one's head (i.e., intuition) instead of the formula (i.e., statistical or mechanical procedure) for clinical prediction. It then describes the controversy that ensued and analyzes the complexity and contemporary relevance of the question itself. Going beyond clinical inference, it identifies select cognitive biases and constraints that cause decision errors, and proposes remedial correctives. Given that the evidence shows cognition to be flawed, the article discusses the linear regression, Bayesian, signal detection, and computer approaches as possible decision aids. Their cost-benefit trade-offs, when used either alone or as complements to one another, are examined and evaluated. The critique concludes with a note of cautious optimism regarding the formula's future role as a decision aid and offers several interim solutions.

Cognition↗

Determination of biogeographical range: an application of molecular phylogeography to the European pool frog Rana lessonae.

Understanding how species are constrained within their biogeographical ranges is a central problem in evolutionary ecology. Essential prerequisites for addressing this question include accurate determinations of range borders and of the genetic structures of component populations. Human translocation of organisms to sites outside their natural range is one factor that increasingly complicates this issue. In areas not far beyond presumed natural range margins it can be particularly difficult to determine whether a species is native or has been introduced. The pool frog (Rana lessonae) in Britain is a specific example of this dilemma . We used variation at six polymorphic microsatellite loci for investigating the phylogeography of R. lessonae and establishing the affinities of specimens from British populations. The existence and distribution of a distinct northern clade of this species in Norway, Sweden and England infer that it is probably a long-standing native of Britain, which should therefore be included within its natural range. This conclusion was further supported by posterior probability estimates using Bayesian clustering. The phylogeographical analysis revealed unexpected patterns of genetic differentiation across the range of R. lessonae that highlighted the importance of historical colonization events in range structuring.

Animals↗

Maximum likelihood haplotyping for general pedigrees.

Haplotype data is valuable in mapping disease-susceptibility genes in the study of Mendelian and complex diseases. We present algorithms for inferring a most likely haplotype configuration for general pedigrees, implemented in the newest version of the genetic linkage analysis system SUPERLINK. In SUPERLINK, genetic linkage analysis problems are represented internally using Bayesian networks. The use of Bayesian networks enables efficient maximum likelihood haplotyping for more complex pedigrees than was previously possible. Furthermore, to support efficient haplotyping for larger pedigrees, we have also incorporated a novel algorithm for determining a better elimination order for the variables of the Bayesian network. The presented optimization algorithm also improves likelihood computations. We present experimental results for the new algorithms on a variety of real and semiartificial data sets, and use our software to evaluate MCMC approximations for haplotyping.

Algorithms↗

Statistical analysis of temporal evolution in single-neuron firing rates.

A fundamental methodology in neurophysiology involves recording the electrical signals associated with individual neurons within brains of awake behaving animals. Traditional statistical analyses have relied mainly on mean firing rates over some epoch (often several hundred milliseconds) that are compared across experimental conditions by analysis of variance. Often, however, the time course of the neuronal firing patterns is of interest, and a more refined procedure can produce substantial additional information. In this paper we compare neuronal firing in the supplementary eye field of a macaque monkey across two experimental conditions. We take the electrical discharges, or 'spikes', to be arrivals in a inhomogeneous Poisson process and then model the firing intensity function using both a simple parametric form and more flexible splines. Our main interest is in making inferences about certain characteristics of the intensity, including the timing of the maximal firing rate. We examine data from 84 neurons individually and also combine results into a hierarchical model. We use Bayesian estimation methods and frequentist significance tests based on a nonparametric bootstrap procedure. We are thereby able to conclude that a substantial fraction of the neurons exhibit important temporal differences in firing intensity across the two conditions, and we quantify the effect across the population of neurons.

Journal Article↗

'O father: where art thou?'--Paternity assessment in an open fission-fusion society of wild bottlenose dolphins (Tursiops sp.) in Shark Bay, Western Australia.

Sexually mature male bottlenose dolphins in Shark Bay cooperate by pursuing distinct alliance strategies to monopolize females in reproductive condition. We present the results of a comprehensive study in a wild cetacean population to test whether male alliance membership is a prerequisite for reproductive success. We compared two methods for inferring paternity: both calculate a likelihood ratio, called the paternity index, between two opposing hypotheses, but they differ in the way that significance is applied to the data. The first method, a Bayesian approach commonly used in human paternity testing, appeared to be overly conservative for our data set, but would be less susceptible to assumptions if a larger number of microsatellite loci had been used. Using the second approach, the computer program cervus 2.0, we successfully assigned 11 paternities to nine males, and 17 paternities to 14 out of 139 sexually mature males at 95% and 80% confidence levels, respectively. It appears that being a member of a bottlenose dolphin alliance is not a prerequisite for paternity: two paternities were obtained by juvenile males (one at the 95%, the other at the 80% confidence level), suggesting that young males without alliance partners pursue different mating tactics to adults. Likelihood analyses showed that these two juvenile males were significantly more likely to be the true father of the offspring than to be their half-sibling (P < 0.05). Using paternity data at an 80% confidence level, we could show that reproductive success was significantly skewed within at least some stable first-order alliances (P < 0.01). Interestingly, there is powerful evidence that one mating was incestuous, with one calf apparently fathered by its mother's father (P < 0.01). Our study suggests that the reproductive success of both allied males, and of nonallied juveniles, needs to be incorporated into an adaptive framework that seeks to explain alliance formation in male bottlenose dolphins.

Animals↗

Phylogenetic relationships among Syndermata inferred from nuclear and mitochondrial gene sequences.

Phylogenetic relationships among Syndermata have been extensively debated, mainly because the sister-group of the Acanthocephala has not yet been clearly identified from analyses of morphological and molecular data. Here we conduct phylogenetic analyses on samples from the 4 classes of Acanthocephala (Archiacanthocephala, Eoacanthocephala, Polyacanthocephala, and Palaeacanthocephala) and the 3 Rotifera classes (Bdelloidea, Monogononta, and Seisonidea). We do so using small-subunit (SSU) and large-subunit (LSU) ribosomal DNA and cytochrome c oxidase subunit 1 (cox 1) sequences. These nuclear and mitochondrial DNA sequences were obtained for 27 acanthocephalans, 9 rotifers, and representatives of 6 phyla that were used as outgroups. Maximum parsimony (MP), maximum likelihood (ML), and Bayesian analyses were conducted on the nuclear rDNA(SSU+LSU) and the combined sequence dataset(SSU+LSU+cox 1 genes). Phylogenetic analyses of the combined rDNA and cox 1 data uniformly provided strong support for a clade including rotifers plus acanthocephalans (Syndermata). Strong support was also found for monophyly of Acanthocephala in analyses of the combined dataset or rDNA sequences alone. Within the Acanthocephala the monophyletic grouping of the representatives of each class was strongly supported. Our results depicted Archiacanthocephala as the sister-group to the remaining acanthocephalans. Analyses of the combined dataset recovered a sister-group relationship between Acanthocephala and Bdelloidea by parsimony, likelihood, and Bayesian methods. Support for this clade was generally strong. Alternative topologies that depicted a different rotifer sister-group of Acanthocephala (or monophyly of Rotifera) were significantly worse. In this paraphyletic assemblage of rotifers, the relative positions of Seisonidea and Monogononta to the clade Bdelloidea+Acanthocephala were inconsistent among trees based on different inference methods. These results indicate that Bdelloidea is the free-living sister-group to acanthocephalans, which should prove key for comparative investigations of the morphological, molecular, and ecological changes accompanying the evolution of parasitism.

Acanthocephala↗

A statistical problem for inference to regulatory structure from associations of gene expression measurements with microarrays.

MOTIVATION: One approach to inferring genetic regulatory structure from microarray measurements of mRNA transcript hybridization is to estimate the associations of gene expression levels measured in repeated samples. The associations may be estimated by correlation coefficients or by conditional frequencies (for discretized measurements) or by some other statistic. Although these procedures have been successfully applied to other areas, their validity when applied to microarray measurements has yet to be tested. RESULTS: This paper describes an elementary statistical difficulty for all such procedures, no matter whether based on Bayesian updating, conditional independence testing, or other machine learning procedures such as simulated annealing or neural net pruning. The difficulty obtains if a number of cells from a common population are aggregated in a measurement of expression levels. Although there are special cases where the conditional associations are preserved under aggregation, in general inference of genetic regulatory structure based on conditional association is unwarranted

Algorithms↗

Bayesian estimation of concordance among gene trees.

Multigene sequence data have great potential for elucidating important and interesting evolutionary processes, but statistical methods for extracting information from such data remain limited. Although various biological processes may cause different genes to have different genealogical histories (and hence different tree topologies), we also may expect that the number of distinct topologies among a set of genes is relatively small compared with the number of possible topologies. Therefore evidence about the tree topology for one gene should influence our inferences of the tree topology on a different gene, but to what extent? In this paper, we present a new approach for modeling and estimating concordance among a set of gene trees given aligned molecular sequence data. Our approach introduces a one-parameter probability distribution to describe the prior distribution of concordance among gene trees. We describe a novel 2-stage Markov chain Monte Carlo (MCMC) method that first obtains independent Bayesian posterior probability distributions for individual genes using standard methods. These posterior distributions are then used as input for a second MCMC procedure that estimates a posterior distribution of gene-to-tree maps (GTMs). The posterior distribution of GTMs can then be summarized to provide revised posterior probability distributions for each gene (taking account of concordance) and to allow estimation of the proportion of the sampled genes for which any given clade is true (the sample-wide concordance factor). Further, under the assumption that the sampled genes are drawn randomly from a genome of known size, we show how one can obtain an estimate, with credibility intervals, on the proportion of the entire genome for which a clade is true (the genome-wide concordance factor). We demonstrate the method on a set of 106 genes from 8 yeast species.

Algorithms↗

Genus Tetrastemma Ehrenberg, 1831 (Phylum Nemertea)--a natural group? Phylogenetic relationships inferred from partial 18S rRNA sequences.

We investigated the monophyletic status of the hoplonemertean taxon Tetrastemma by reconstructing the phylogeny for 22 specimens assigned to this genus, together with another 25 specimens from closely related hoplonemertean genera. The phylogeny was based on partial 18S rRNA sequences using Bayesian and maximum likelihood analyses. The included Tetrastemma-species formed a well-supported clade, although the within-taxon relationships were unsettled. We conclude that the name Tetrastemma refers to a monophyletic taxon, but that it cannot be defined by morphological synapomorphies, and our results do not imply that all the over 100 species assigned to this genus belong to it. The results furthermore indicate that the genera Amphiporus and Emplectonema are non-monophyletic.

Animals↗

Phylogeography of Australia's king brown snake (Pseudechis australis) reveals Pliocene divergence and Pleistocene dispersal of a top predator.

King brown snakes or mulga snakes (Pseudechis australis) are the largest and among the most dangerous and wide-ranging venomous snakes in Australia and New Guinea. They occur in diverse habitats, are important predators, and exhibit considerable morphological variation. We infer the relationships and historical biogeography of P. australis based on phylogenetic analysis of 1,249 base pairs from the mitochondrial cytochrome b, NADH dehydrogenase subunit 4 and three adjacent tRNA genes using Bayesian, maximum-likelihood, and maximum-parsimony methods. All methods reveal deep phylogenetic structure with four strongly supported clades comprising snakes from New Guinea (I), localities all over Australia (II), the Kimberleys of Western Australia (III), and north-central Australia (IV), suggesting a much more ancient radiation than previously believed. This conclusion is robust to different molecular clock estimations indicating divergence in Pliocene or Late Miocene, after landbridge dispersal to New Guinea had occurred. While members of clades I, III and IV are medium-sized, slender snakes, those of clade II attain large sizes and a robust build, rendering them top predators in their ecosystems. Genetic differentiation within clade II is low and haplotype distribution largely incongruent with geography or colour morphs, suggesting Pleistocene dispersal and recent ecomorph evolution. Significant haplotype diversity exists in clades III and IV, implying that clade IV comprises two species. Members of clade II are broadly sympatric with members of both northern Australian clades. Thus, our data support the recognition of at least five species from within P. australis (auct.) under various criteria. We discuss biogeographical, ecological and medical implications of our findings.

Animals↗

Evolution of Piperales--matK gene and trnK intron sequence data reveal lineage specific resolution contrast.

Piperales represent the largest basal angiosperm order with a nearly worldwide distribution. The order includes three species rich genera, Piper (ca. 2000 species), Peperomia (ca. 1500-1700 species), and Aristolochia s. l. (ca. 500 species). Sequences of the matK gene and the non-coding trnK group II intron are analysed for a dense set of 105 taxa representing all families (except Hydnoraceae) and all generic segregates (except Euglypha within Aristolochiaceae) of Piperales. A large number of highly informative indels are found in the Piperales trnK/matK dataset. Within a narrow region approximately 500 nt downstream in the matK coding region (CDS), a length variable simple sequence repeat (SSR) expansion segment occurs, in which insertions and deletions have led to short frame-shifts. These are corrected shortly afterwards, resulting in a maximum of six amino acids being affected. Furthermore, additional non-functional matK copies were found in Zippelia begoniifolia, which can easily be discriminated from the functional open reading frame (ORF). The trnK/matK sequence data fully resolve relationships within Peperomia, whereas they are not effective within Piper. The resolution contrast is correlated with the rate heterogeneity between those lineages. Parsimony, Bayesian and likelihood analyses result in virtually the same topology, and converge on the monophyly of Piperaceae and Saururaceae. Lactoris gains high support as sister to Aristolochiaceae subf. Aristolochioideae, but the different tree inference methods yield conflicting results with respect to the relationships of subfam. Asaroideae. In Piperaceae, a clade formed by the monotypic genus Zippelia and the small genus Manekia (=Sarcorhachis) is sister to the two large genera Piper and Peperomia.

Base Sequence↗

A Thurstonian model for quantitative genetic analysis of ranks: a Bayesian approach.

A fully Bayesian method for quantitative genetic analysis of data consisting of ranks of, e.g., genotypes, scored at a series of events or experiments is presented. The model postulates a latent structure, with an underlying variable realized for each genotype or individual involved in the event. The rank observed is assumed to reflect the order of the values of the unobserved variables, i.e., the classical Thurstonian model of psychometrics. Parameters driving the Bayesian hierarchical model include effects of covariates, additive genetic effects, permanent environmental deviations, and components of variance. A Markov chain Monte Carlo implementation based on the Gibbs sampler is described, and procedures for inferring the probability of yet to be observed future rankings are outlined. Part of the model is rendered nonparametric by introducing a Dirichlet process prior for the distribution of permanent environmental effects. This can lead to potential identification of clusters of such effects, which, in some competitions such as horse races, may reflect forms of undeclared preferential treatment.

Bayes Theorem↗