Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Comprehensive comparative analysis of kinesins in photosynthetic eukaryotes.

BACKGROUND: Kinesins, a superfamily of molecular motors, use microtubules as tracks and transport diverse cellular cargoes. All kinesins contain a highly conserved approximately 350 amino acid motor domain. Previous analysis of the completed genome sequence of one flowering plant (Arabidopsis) has resulted in identification of 61 kinesins. The recent completion of genome sequencing of several photosynthetic and non-photosynthetic eukaryotes that belong to divergent lineages offers a unique opportunity to conduct a comprehensive comparative analysis of kinesins in plant and non-plant systems and infer their evolutionary relationships. RESULTS: We used the kinesin motor domain to identify kinesins in the completed genome sequences of 19 species, including 13 newly sequenced genomes. Among the newly analyzed genomes, six represent photosynthetic eukaryotes. A total of 529 kinesins was used to perform comprehensive analysis of kinesins and to construct gene trees using the Bayesian and parsimony approaches. The previously recognized 14 families of kinesins are resolved as distinct lineages in our inferred gene tree. At least three of the 14 kinesin families are not represented in flowering plants. Chlamydomonas, a green alga that is part of the lineage that includes land plants, has at least nine of the 14 known kinesin families. Seven of ten families present in flowering plants are represented in Chlamydomonas, indicating that these families were retained in both the flowering-plant and green algae lineages. CONCLUSION: The increase in the number of kinesins in flowering plants is due to vast expansion of the Kinesin-14 and Kinesin-7 families. The Kinesin-14 family, which typically contains a C-terminal motor, has many plant kinesins that have the motor domain at the N terminus, in the middle, or the C terminus. Several domains in kinesins are present exclusively either in plant or animal lineages. Addition of novel domains to kinesins in lineage-specific groups contributed to the functional diversification of kinesins. Results from our gene-tree analyses indicate that there was tremendous lineage-specific duplication and diversification of kinesins in eukaryotes. Since the functions of only a few plant kinesins are reported in the literature, this comprehensive comparative analysis will be useful in designing functional studies with photosynthetic eukaryotes.

Algal Proteins↗

Model selection and model averaging in phylogenetics: advantages of akaike information criterion and bayesian approaches over likelihood ratio tests.

Model selection is a topic of special relevance in molecular phylogenetics that affects many, if not all, stages of phylogenetic inference. Here we discuss some fundamental concepts and techniques of model selection in the context of phylogenetics. We start by reviewing different aspects of the selection of substitution models in phylogenetics from a theoretical, philosophical and practical point of view, and summarize this comparison in table format. We argue that the most commonly implemented model selection approach, the hierarchical likelihood ratio test, is not the optimal strategy for model selection in phylogenetics, and that approaches like the Akaike Information Criterion (AIC) and Bayesian methods offer important advantages. In particular, the latter two methods are able to simultaneously compare multiple nested or nonnested models, assess model selection uncertainty, and allow for the estimation of phylogenies and model parameters using all available models (model-averaged inference or multimodel inference). We also describe how the relative importance of the different parameters included in substitution models can be depicted. To illustrate some of these points, we have applied AIC-based model averaging to 37 mitochondrial DNA sequences from the subgenus Ohomopterus(genus Carabus) ground beetles described by Sota and Vogler (2001).

Animals↗

Global diversity and evolution of Salmonella enterica serovar Panama: a genomic epidemiology study.

BACKGROUND: Non-typhoidal Salmonella is a globally important bacterial pathogen, typically associated with foodborne gastrointestinal infection. Some non-typhoidal Salmonella serovars can also colonise typically sterile sites in people to cause invasive non-typhoidal Salmonella disease. Salmonella enterica serovar Panama is responsible for a substantial number of cases of human bloodstream infection, but despite its global dissemination, numerous outbreaks, and a reported association with invasive non-typhoidal Salmonella disease, S enterica serovar Panama (S Panama) is understudied. We aimed to describe the genomic epidemiology and evolutionary history of S Panama to provide a vital baseline of understanding for this globally important serovar. METHODS: In this genomic epidemiology study, we analysed S Panama genomes derived from historical collections, national surveillance datasets, and publicly available epidemiological and whole-genome sequencing data which span the years 1931-2019. Maximum likelihood and Bayesian phylodynamic approaches were used to investigate population structure and evolutionary history and to infer geotemporal dissemination. A combination of different bioinformatic approaches with short-read and long-read data were used to characterise geographical and clade-specific trends in antimicrobial resistance (AMR) and genetic markers for invasiveness. FINDINGS: We analysed 836 S Panama genomes, of which 559 (67%) were sequenced as part of this study. The collection represents all inhabited continents and includes isolates collected between 1931 and 2019. We identified the presence of four geographically linked S Panama clades (C1 [ie, the Latin America and the Caribbean clade; n=338], C2 [ie, the European clade; n=124], C3 [ie, the Martinique clade; n=131], and C4 [ie, the Asia and Oceania clade; n=104]) and regional trends in AMR profiles. Most isolates (715 [86%] of 836) were pan-susceptible to antibiotics and belonged to clades circulating in Latin America and the Caribbean (64%, n=458). Most antibiotic-resistant isolates in our collection (113 [93%] of 121) fell within clades C4 (ie, the Asia and Oceania clade) and C2 (ie, the European clade), the latter of which had the highest invasiveness index values based on the conservation of 196 extraintestinal predictor genes. INTERPRETATION: This first large-scale phylogenetic analysis of S Panama has revealed important information about the population structure, AMR, global ecology, and genetic markers of invasiveness of the identified genomic subtypes. Our findings provide an important baseline for understanding S Panama infection. The presence of multidrug-resistant clades with elevated invasiveness index values should be monitored through ongoing surveillance, as such clades could pose an increased public health risk. FUNDING: UK Research and Innovation Global Challenges Research Fund and Biotechnology and Biological Sciences Research Council, UK Medical Research Council, Wellcome Trust, John Lennon Memorial Scholarship, Institut Pasteur, Santé publique France, Fondation Le Roch-Les Mousquetaires, Investissement d'Avenir Programme, and Australian National Health and Medical Research Council.

Humans↗

Molecular systematics of Zopfiella and allied genera: evidence from multi-gene sequence analyses.

This study aims to reveal the phylogenetic relationships of Zopfiella and allied genera in the Sordariales. Multiple gene sequences (partial 28S rDNA, ITS/5.8S rDNA and partial beta-tubulin) were analysed using MP and Bayesian analyses. Analyses of different gene datasets were performed individually and then combined to infer phylogenies. Phylogenetic analyses show that currently recognised Zopfiella species are polyphyletic. Based on sequence analyses and morphology, it appears that Zopfiella should be restricted to species having ascospores with a septum in the dark cell. Our molecular analysis also shows that Zopfiella should be placed in Lasiosphaeriaceae rather than Chaetomiaceae. Cercophora and Podospora are also polyphyletic, which is in agreement with previous studies. Our analyses show that species possessing a Cladorrhinum anamorph are phylogenetically closely related. In addition, there are several strongly supported clades, characterised by species possessing divergent morphological characters. It is difficult to predict which characters are phylogenetically informative for delimiting these clades.

Base Sequence↗

Using Bayesian statistics to estimate the coefficients of a two- component second-order chlorine bulk decay model for a water distribution system.

Most chlorine decay models for the bulk phase in a water distribution system consider only chlorine concentration and time. Clark [1998. Chlorine demand and trihalomethane formation kinetics: a second-order model. J. Environ. Eng. 124(1), 16-24] first proposed a two-component second-order chlorine decay model based on the concept of competing reacting substances. A corrected mathematical formulation is developed and, because the recent findings suggested that not all natural organic matter (NOM) is involved in the chlorine decay process, an additional parameter is introduced. A parameter assignment method employing Bayesian statistical analysis incorporating Monte Carlo Markov chain (MCMC) with Gibbs sampling to make inferences, is employed in the estimation of model parameters. Three parameters are estimated for the model, namely the ratio of chlorine to TOC, the chlorine reaction rate, and a fraction factor of TOC which represents the true amount of TOC involved in chlorine decay process. Water samples taken from Goderich in the summer of 2005, are used for estimating the parameters.

Bayes Theorem↗

Modeling of trough plasma bismuth concentrations.

Disposition pharmacokinetics of bismuth following oral dosing of ranitidine bismuth citrate are complicated and variable. An analysis of data from healthy volunteers suggests a model with three disposition compartments and first-order absorption. Patient data are pooled from 10 separate studies and consist of 1140 trough concentrations measured in 802 patients following dosing of 2 to 12 weeks duration. There are therefore insufficient data to obtain reliable parameter estimates for the full model and we use instead a much reduced model and an informative prior based on the volunteer data. Individual parameter estimates from this model can then be used to establish covariate relationships. Trough concentrations were influenced by the coadministration of clarithromycin and by creatinine clearance. A simulation study was carried out to check the validity of the estimates obtained from the reduced model. We carry out analysis via Bayesian sampling-based techniques. Throughout, we use predictive distributions for both diagnostic and inference purposes. In particular, we determine predicted distributions for the Cmax, Cmin and AUC characteristics of new individuals.

Bayes Theorem↗

Testing species boundaries in an ancient species complex with deep phylogeographic history: genus Xantusia (Squamata: Xantusiidae).

Identification of species in natural populations has recently received increased attention with a number of investigators proposing rigorous methods for species delimitation. Morphologically conservative species (or species complexes) with deep phylogenetic histories (and limited gene flow) are likely to pose particular problems when attempting to delimit species, yet this is crucial to comparative studies of the geography of speciation. We apply two methods of species delimitation to an ancient group of lizards (genus Xantusia) that occur throughout southwestern North America. Mitochondrial cytochrome b and nicotinamide adenine dinucleotide dehydrogenase subunit 4 gene sequences were generated from samples taken throughout the geographic range of Xantusia. Maximum likelihood, Bayesian, and nested cladogram analyses were used to estimate relationships among haplotypes and to infer evolutionary processes. We found multiple well-supported independent lineages within Xantusia, for which there is considerable discordance with the currently recognized taxonomy. High levels of sequence divergence (21.3%) suggest that the pattern in Xantusia may predate the vicariant events usually hypothesized for the fauna of the Baja California peninsula, and the existence of deeply divergent clades (18.8%-26.9%) elsewhere in the complex indicates the occurrence of ancient sundering events whose genetic signatures were not erased by the late Wisconsin vegetation changes. We present a revised taxonomic arrangement for this genus consistent with the distinct mtDNA lineages and discuss the phylogeographic history of this genus as a model system for studies of speciation in North American deserts.

Animals↗

Automated semantic analysis of changes in image sequences of neurons in culture.

Quantitative studies of dynamic behaviors of live neurons are currently limited by the slowness, subjectivity, and tedium of manual analysis of changes in time-lapse image sequences. Challenges to automation include the complexity of the changes of interest, the presence of obfuscating and uninteresting changes due to illumination variations and other imaging artifacts, and the sheer volume of recorded data. This paper describes a highly automated approach that not only detects the interesting changes selectively, but also generates quantitative analyses at multiple levels of detail. Detailed quantitative neuronal morphometry is generated for each frame. Frame-to-frame neuronal changes are measured and labeled as growth, shrinkage, merging, or splitting, as would be done by a human expert. Finally, events unfolding over longer durations, such as apoptosis and axonal specification, are automatically inferred from the short-term changes. The proposed method is based on a Bayesian model selection criterion that leverages a set of short-term neurite change models and takes into account additional evidence provided by an illumination-insensitive change mask. An automated neuron tracing algorithm is used to identify the objects of interest in each frame. A novel curve distance measure and weighted bipartite graph matching are used to compare and associate neurites in successive frames. A separate set of multi-image change models drives the identification of longer term events. The method achieved frame-to-frame change labeling accuracies ranging from 85% to 100% when tested on 8 representative recordings performed under varied imaging and culturing conditions, and successfully detected all higher order events of interest. Two sequences were used for training the models and tuning their parameters; the learned parameter settings can be applied to hundreds of similar image sequences, provided imaging and culturing conditions are similar to the training set. The proposed approach is a substantial innovation over manual annotation and change analysis, accomplishing in minutes what it would take an expert hours to complete.

Algorithms↗

Analyzing gene expression time-courses.

Measuring gene expression over time can provide important insights into basic cellular processes. Identifying groups of genes with similar expression time-courses is a crucial first step in the analysis. As biologically relevant groups frequently overlap, due to genes having several distinct roles in those cellular processes, this is a difficult problem for classical clustering methods. We use a mixture model to circumvent this principal problem, with hidden Markov models (HMMs) as effective and flexible components. We show that the ensuing estimation problem can be addressed with additional labeled data-partially supervised learning of mixtures-through a modification of the Expectation-Maximization (EM) algorithm. Good starting points for the mixture estimation are obtained through a modification to Bayesian model merging, which allows us to learn a collection of initial HMMs. We infer groups from mixtures with a simple information-theoretic decoding heuristic, which quantifies the level of ambiguity in group assignment. The effectiveness is shown with high-quality annotation data. As the HMMs we propose capture asynchronous behavior by design, the groups we find are also asynchronous. Synchronous subgroups are obtained from a novel algorithm based on Viterbi paths. We show the suitability of our HMM mixture approach on biological and simulated data and through the favorable comparison with previous approaches. A software implementing the method is freely available under the GPL from http://ghmm.org/gql.

Algorithms↗

Robustness considerations in Bayesian analysis.

All statistical analyses demand uncertain inputs or assumptions. This is especially true of Bayesian analyses. In addition to the usual concerns about the agreement of the data and model, a Bayesian must contemplate the effect of an uncertain prior specification. The degree to which inferences are robust to changes in the prior is of primary interest. This article discusses some robust techniques that have been suggested in the literature. One goal is to make apparent the relevance of some of these techniques to biostatistical work.

Adrenergic beta-Antagonists↗

A four-shock Bayesian up-down estimator of the 80% effective defibrillation dose.

INTRODUCTION: New defibrillation techniques are often compared to standard approaches using the defibrillation threshold. However, inference from thresholding data necessitates extrapolation from reactions to relatively ineffective shocks, an error prone procedure requiring large sample sizes for hypothesis testing and large safety margins for defibrillator implantation. In contrast, this article presents a clinically validated statistical model of a minimum error, four-shock defibrillation testing protocol for estimating the 80% effective defibrillation strength for a given patient (ED80). METHODS AND RESULTS: A Bayesian statistical model was constructed assuming that the defibrillation dose-response curve is sigmoidal, and the ED80 is between 150 and 750 V. The model was used to design a minimum predicted error testing protocol and estimates. To prospectively validate the testing protocol and estimates, 170 patients received voltage-programmed biphasic testing. Four fibrillation episodes were induced and terminated in each patient according to the Bayesian up-down protocol. In addition, a validation attempt was made at the estimated ED80 rounded up to the nearest 50 V. In order to estimate the safety margin, in 136 patients, a defibrillation attempt was made at the rounded ED80 + 100 V. Of the 170 attempts at the rounded ED80, 143 (84%) attempts terminated fibrillation. Of the 136 attempts at the rounded ED80 + 100 V, 133 (98%) were effective. CONCLUSIONS: The four-shock Bayesian up-down protocol is the first clinical protocol to accurately predict an ED80 voltage. A 100 V increment above the ED80 provides an adequate safety margin. This simple and accurate method for estimating a highly effective defibrillation dose may be a valuable tool for population-based clinical hypothesis testing, as well as defibrillator implantation.

Adult↗

Bayesian estimation of parameters of a structural model for genetic covariances between milk yield in five regions of the United States.

Inference about genetic covariance matrices using multiple-trait models is often hindered by lack of information. This leads to imprecise estimates of genetic parameters and of breeding values. Patterns in a genetic covariance matrix can be exploited to reduce the number of parameters and to increase quality of inferences. A structural model for genetic covariances was developed and fitted to milk yield data in five regions of the United States. This was compared with a standard multiple-trait analysis using a deviance information criterion, a measure of quality of fit. Data consisted of 3,465,334 Holstein first-lactation records from daughters of 43,755 sires in five regions of the United States (Midwest, Northeast, Northwest, Southeast, Southwest). Parameters of the structural model included an intercept and effects of measures of genetic and of management similarity on genetic covariances. Genetic similarity depended on the number of records contributed by sires that were common to a pair of regions. Management similarity was a function of the quantity of concentrate used to produce 1000 kg of milk in each pair of regions. The structural and the multiple-trait models gave similar estimates of genetic covariances, but the number of parameters was 8 in the former vs. 15 in the latter. Hence, estimates of genetic covariances were more precise with the structural model. A deviance information criterion suggested a slight superiority of the multiple-trait model, although probably within sampling error. For both models, genetic correlations between milk yield in five regions of the United States were larger than 0.93.

Analysis of Variance↗

ARACNE: an algorithm for the reconstruction of gene regulatory networks in a mammalian cellular context.

BACKGROUND: Elucidating gene regulatory networks is crucial for understanding normal cell physiology and complex pathologic phenotypes. Existing computational methods for the genome-wide "reverse engineering" of such networks have been successful only for lower eukaryotes with simple genomes. Here we present ARACNE, a novel algorithm, using microarray expression profiles, specifically designed to scale up to the complexity of regulatory networks in mammalian cells, yet general enough to address a wider range of network deconvolution problems. This method uses an information theoretic approach to eliminate the majority of indirect interactions inferred by co-expression methods. RESULTS: We prove that ARACNE reconstructs the network exactly (asymptotically) if the effect of loops in the network topology is negligible, and we show that the algorithm works well in practice, even in the presence of numerous loops and complex topologies. We assess ARACNE's ability to reconstruct transcriptional regulatory networks using both a realistic synthetic dataset and a microarray dataset from human B cells. On synthetic datasets ARACNE achieves very low error rates and outperforms established methods, such as Relevance Networks and Bayesian Networks. Application to the deconvolution of genetic networks in human B cells demonstrates ARACNE's ability to infer validated transcriptional targets of the cMYC proto-oncogene. We also study the effects of misestimation of mutual information on network reconstruction, and show that algorithms based on mutual information ranking are more resilient to estimation errors. CONCLUSION: ARACNE shows promise in identifying direct transcriptional interactions in mammalian cellular networks, a problem that has challenged existing reverse engineering algorithms. This approach should enhance our ability to use microarray data to elucidate functional mechanisms that underlie cellular processes and to identify molecular targets of pharmacological compounds in mammalian cellular networks.

Algorithms↗

Phylogeny of diplozoids in five genera of the subfamily Diplozoinae Palombi, 1949 as inferred from ITS-2 rDNA sequences.

The phylogenetic relationship of 5 genera, i.e. Diplozoon Nordmann, 1832, Paradiplozoon Achmerov, 1974, Inustiatus Khotenovsky, 1978, Sindiplozoon Khotenovsky, 1981, and Eudiplozoon Khotenovsky, 1985 in the subfamily Diplozoinae Palombi, 1949 (Monogenea, Polyopisthocotylea) was inferred from rDNA ITS-2 region using neighbour-joining (NJ), maximum likelihood (ML) and Bayesian methods. The phylogenetic trees produced by using NJ, ML and Bayesian methods exhibit essentially the same topology. Surprisingly, freshwater species of Paradiplozoon from Europe clustered together with species of Diplozoon, but separated from Chinese Paradiplozoon species. The results of molecular phylogeny and lower level of divergence (4.1-15.7%) in ITS-2 rDNA among Paradiplozoon from Europe and Diplozoon and, on the other hand, high level of divergence (45.3-53.7%) among Paradiplozoon species from Europe and China might indicate the non-monophyletic origin of the genus Paradiplozoon. Also, the generic status of European Paradiplozoon needs to be revised. The species of Paradiplozoon in China is a basal group in Diplozoinae as revealed by NJ and Bayesian methods, and Sindiplozoon appears to be closely related to European Paradiplozoon and Diplozoon with their relationship to Eudiplozoon and Inustiatus being unresolved.

Animals↗

Genome-wide analysis of mouse transcripts using exon microarrays and factor graphs.

Recent mammalian microarray experiments detected widespread transcription and indicated that there may be many undiscovered multiple-exon protein-coding genes. To explore this possibility, we labeled cDNA from unamplified, polyadenylation-selected RNA samples from 37 mouse tissues to microarrays encompassing 1.14 million exon probes. We analyzed these data using GenRate, a Bayesian algorithm that uses a genome-wide scoring function in a factor graph to infer genes. At a stringent exon false detection rate of 2.7%, GenRate detected 12,145 gene-length transcripts and confirmed 81% of the 10,000 most highly expressed known genes. Notably, our analysis showed that most of the 155,839 exons detected by GenRate were associated with known genes, providing microarray-based evidence that most multiple-exon genes have already been identified. GenRate also detected tens of thousands of potential new exons and reconciled discrepancies in current cDNA databases by 'stitching' new transcribed regions into previously annotated genes.

Algorithms↗

Finding scientific topics.

A first step in identifying the content of a document is determining which topics that document addresses. We describe a generative model for documents, introduced by Blei, Ng, and Jordan [Blei, D. M., Ng, A. Y. & Jordan, M. I. (2003) J. Machine Learn. Res. 3, 993-1022], in which each document is generated by choosing a distribution over topics and then choosing each word in the document from a topic selected according to this distribution. We then present a Markov chain Monte Carlo algorithm for inference in this model. We use this algorithm to analyze abstracts from PNAS by using Bayesian model selection to establish the number of topics. We show that the extracted topics capture meaningful structure in the data, consistent with the class designations provided by the authors of the articles, and outline further applications of this analysis, including identifying "hot topics" by examining temporal dynamics and tagging abstracts to illustrate semantic content.

Databases, Factual↗

The complete mitochondrial genome of the rayfish Raja porosa (Chondrichthyes, Rajidae).

We isolated mitochondrial DNA from the rayfish Raja porosa by long-polymerase chain reaction (Long-PCR) with conserved primers, and sequenced it by primer walking method using flanking sequences as sequencing primers. R. porosa mitochondrial DNA consists of 16,972 bp and its structural organization is conserved in comparison with other fishes and mammals. Based on the mitochondrial cytochrome b (cyt b) sequence, the phylogenetic position of R. porosa among cartilaginous fishes was inferred using different phylogenetic methods (ML-based quartet puzzling, Neighbor-joining (NJ) and Bayesian approaches). In this paper, we report the characteristics of the R. porosa mitochondrial genome including structural organization, base composition of rRNAs, tRNAs and protein-encoding genes and characteristics of mitochondrial tRNAs. These findings are applicable to comparative mitogenomics of R. porosa with other related taxa.

Amino Acid Motifs↗

VAMPIRE microarray suite: a web-based platform for the interpretation of gene expression data.

Microarrays are invaluable high-throughput tools used to snapshot the gene expression profiles of cells and tissues. Among the most basic and fundamental questions asked of microarray data is whether individual genes are significantly activated or repressed by a particular stimulus. We have previously presented two Bayesian statistical methods for this level of analysis, collectively known as variance-modeled posterior inference with regional exponentials (VAMPIRE). These methods each require a sophisticated modeling step followed by integration of a posterior probability density. We present here a publicly available, web-based platform that allows users to easily load data, associate related samples and identify differentially expressed features using the VAMPIRE statistical framework. In addition, this suite of tools seamlessly integrates a novel gene annotation tool, known as GOby, which identifies statistically overrepresented gene groups. Unlike other tools in this genre, GOby can localize enrichment while respecting the hierarchical structure of annotation systems like Gene Ontology (GO). By identifying statistically significant enrichment of GO terms, Kyoto Encyclopedia of Genes and Genomes pathways, and TRANSFAC transcription factor binding sites, users can gain substantial insight into the physiological significance of sets of differentially expressed genes. The VAMPIRE microarray suite can be accessed at http://genome.ucsd.edu/microarray.

Bayes Theorem↗