Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Phylogenetic relationships of Peronospora and related genera based on nuclear ribosomal ITS sequences.

In order to investigate phylogenetic relationships of selected members of Peronosporaceae to the genera Pythium, Halophkytophthora, Phytophthora, and Peronophythora, Bayesian analysis of partial sequences of the ITS1-5.8S-ITS2 region was performed. In addition, sequences of the complete ITS1-5.8S-ITS2 region were analysed for 101 collections belonging to the genera Peronospora, Hyaloperonospora, Perofascia, Pseudoperonospora, Phytophthora and Peronophythora, using Bayesian inference of phylogeny and maximum parsimony. The results confirm a close relationship of these genera. The strongly supported Peronosporaceae clade is located within the paraphyletic genus Phytophthora (including Peronophythora). Monophyly of the genera Pseudoperonospora and Hyaloperonospora are strongly supported, but monophyly of Peronospora s. str. can neither be confirmed nor rejected. Within Peronospora s. str., basic relationships often remain unclear; however, some groups form highly supported monophyletic clades. Peronospora species parasitising the same host families or orders are only partly resolved as monophyletic, indicating frequent host-jumping also between distantly related host families. All species inhabiting flowers of different host families form a strongly supported monophyletic group.

Base Sequence↗

Molecular phylogeny of Ustilago, Sporisorium, and related taxa based on combined analyses of rDNA sequences.

Combined analyses of ITS and LSU rDNA sequences were utilized to resolve the phylogenetic relationships of 98 members of the smut genera Lundquistia, Melanopsichium, Moesziomyces, Macalpinomyces, Sporisorium, and Ustilago (Basidiomycota: Ustilaginales). Minimum Evolution and Bayesian inference of phylogeny resolve three major groups of almost identical composition: Sporisorium, Ustilago, and a basal assemblage of both Ustilago and Sporisorium species. Macalpinomyces deserves generic rank regarding its type species M. eriachnes; all other Macalpinomyces species of our study clearly turn out to be part of Ustilago or Sporisorium. Lundquistia evidently belongs to Sporisorium. Moesziomyces, probably paraphyletic, stands basal to all other genera. Interestingly, Melanopsichium belongs to the Ustilago clade, being the only member of the ingroup not parasitizing on Poaceae. The patchy distribution of commonly used morphological characters along our phylograms points to their variability and dependence on the host's morphological traits instead of being valuable for resolving parasite phylogeny. The new combination: Sporisorium fascicularis comb. nov. (syn. Lundquistia.fascicularis) is made.

DNA, Fungal↗

A re-consideration of Pseudoperonospora cubensis and P. humuli based on molecular and morphological data.

Phylogenetic analysis of the ITS rDNA region was carried out with two economically important downy mildews, Pseudoperonospora cubensis, which infects species of Cucumis, Cucurbita, and Citrullus belonging to Cucurbitaceae, and P. humuli, which infects plants of the genus Humulus belonging to Cannabaceae. Two closely related species, P. cannabina and P. celtidis, were also included to reveal taxonomic relationships with the first two mildews. All four species formed a well-resolved clade when compared with the ITS sequences of other downy mildew genera, using Bayesian inference and maximum parsimony. The P. cubensis isolates obtained from different hosts and (or) geographical origins in Korea, exhibited no intraspecific variability in the ITS sequences. The phylogenetic analyses of P. cubensis and P. humuli showed that they share a high level of sequence homology; the morphology of the sporangiophores, sporangia, and dehiscence apparatus confirmed the similarity of the two species. We therefore reduce P. humuli to the status of a taxonomic synonym of P. cubensis.

DNA, Fungal↗

A comparison of frailty and other models for bivariate survival data.

Multivariate survival data arise when each study subject may experience multiple events or when study subjects are clustered into groups. Statistical analyses of such data need to account for the intra-cluster dependence through appropriate modeling. Frailty models are the most popular for such failure time data. However, there are other approaches which model the dependence structure directly. In this article, we compare the frailty models for bivariate data with the models based on bivariate exponential and Weibull distributions. Bayesian methods provide a convenient paradigm for comparing the two sets of models we consider. Our techniques are illustrated using two examples. One simulated example demonstrates model choice methods developed in this paper and the other example, based on a practical data set of onset of blindness among patients with diabetic Retinopathy, considers Bayesian inference using different models.

Bayes Theorem↗

Assessing the impact of managed-care on the distribution of length-of-stay using Bayesian hierarchical models.

Hierarchical models provide a useful framework for the complexities encountered in policy-relevant research in which the impact of social programs is being assessed. Such complexities include multi-site data, censored data and over-dispersion. In this paper, Bayesian inference through Markov Chain Monte Carlo methods is used for the analysis of a complex hierarchical log-normal model that shows the impact of a managed care strategy aimed at limiting length of hospital stays. Parameters in this model allow for variability in baseline length-of-stay as well as the program effect across hospitals. The authors demonstrate elicitation and sensitivity analysis with respect to prior distributions. All calculations for the posterior and predictive distributions were obtained using the software BUGS.

Bayes Theorem↗

Bayesian methods for missing covariates in cure rate models.

We propose methods for Bayesian inference for missing covariate data with a novel class of semiparametric survival models with a cure fraction. We allow the missing covariates to be either categorical or continuous and specify a parametric distribution for the covariates that is written as a sequence of one dimensional conditional distributions. We assume that the missing covariates are missing at random (MAR) throughout. We propose an informative class of joint prior distributions for the regression coefficients and the parameters arising from the covariate distributions. The proposed class of priors are shown to be useful in recovering information on the missing covariates especially in situations where the missing data fraction is large. Properties of the proposed prior and resulting posterior distributions are examined. Also, model checking techniques are proposed for sensitivity analyses and for checking the goodness of fit of a particular model. Specifically, we extend the Conditional Predictive Ordinate (CPO) statistic to assess goodness of fit in the presence of missing covariate data. Computational techniques using the Gibbs sampler are implemented. A real data set involving a melanoma cancer clinical trial is examined to demonstrate the methodology.

Bayes Theorem↗

Sorting signals from protein NMR spectra: SPI, a Bayesian protocol for uncovering spin systems.

Grouping of spectral peaks into J-connected spin systems is essential in the analysis of macromolecular NMR data as it provides the basis for disentangling chemical shift degeneracies. It is a mandatory step before resonance and NOESY cross-peak identities can be established. We have developed SPI, a computational protocol that scrutinizes peak lists from homo- and hetero-nuclear multidimensional NMR spectra and progressively assembles sets of resonances into consensus J- and/or NOE-connected spin systems. SPI estimates the likelihood of nuclear spin resonances appearing at defined frequencies given sets of cross-peaks measured from multi-dimensional experiments. It quantifies spin system matching probabilities via Bayesian inference. The protocol takes advantage of redundancies in the number of connectivities revealed by suites of diverse NMR experiments, systematically tracking the adequacy of each grouping hypothesis. SPI was tested on 2D homonuclear and 2D/3D(15)N-edited data recorded from two protein modules, the col 2 domain of matrix metalloproteinase-2 (MMP-2) and the kringle 2 domain of plasminogen, of 60 and 83 amino acid residues, respectively. For these protein domains SPI identifies approximately 95% unambiguous resonance frequencies, a relatively good performance vis-à-vis the reported 'manual' (interactive) analyses. Abbreviations and Acronyms: SPI, SPin Identification; BMRB, BioMagResBank (Madison, WI).

Amino Acid Sequence↗

Physiologically-based pharmacokinetics and molecular pharmacodynamics of 17-(allylamino)-17-demethoxygeldanamycin and its active metabolite in tumor-bearing mice.

A whole-body physiologically-based model was developed to describe the pharmacokinetics of the ansamycin benzoquinone antibiotic 17-(allylamino)-17-demethoxygeldanamycin (17AAG) and its active metabolite 17-(amino-)-17-demethoxygeldanamycin (17AG) in blood, normal organs (lung, brain, heart, spleen, liver, kidney, skeletal muscle) and implanted human tumor xenograft in nude mice. The distribution of 17 AAG in all organs was described by diffusion-limited exchange models, while that of 17 AG was described by perfusion-limited models. The intrinsic clearances of 17AAG and 17AG in the liver were uniquely identified using local models and were estimated to be 4.93 ml/hr and 3.34 ml/hr. It was also estimated that the formation of 17AG in liver accounted for 40% of the 17AAG intrinsic clearance. The model for the distribution of both 17AAG and 17AG in the human breast cancer tumor xenograft included vascular, interstitial and intracellular compartments, which yielded the predicted cellular concentrations of 17AAG and 17AG two to three times higher than the corresponding whole tissue measurements at steady state. Estimates of the vascular-interstitial permeability surface-area product were similar for 17AAG and 17AG (0.23 ml/hr and 0.26 ml/hr). However, the interstitial to cellular transport rate of 17AG was three-fold greater than that of 17AAG, which resulted in the preferential uptake of 17AG over 17AAG in tumor. Indirect response models were developed to describe the combined action of 17AAG and 17AG on the onco-proteins Raf-1 and p185erbB2 in tumor. The half-life of endogenous protein turnover was estimated to be 22.6 hr for Raf-1 and 8.6 hr for p185erbB2, and both were comparable to corresponding values measured in vitro. A model for the molecular chaperon heat shock proteins HSP70 and HSP90 was developed based on the molecular mechanism of heat shock auto-regulation and the action of 17AAG and 17AG on these proteins. The model provided in vivo estimates of endogenous HSP70 and HSP90 turnover. In modeling pharmacokmetics and pharmacodynamics, Bayesian inference was employed to estimate the kinetic, physiological and molecular parameters when prior information was available.

Animals↗

Molecular phylogenetic analysis of the Microphalloidea Ward, 1901 (Trematoda: Digenea).

Phylogenetic interrelationships of 32 species belonging to 18 genera and four families of the superfamily Microphalloidea were studied using partial sequences of nuclear lsrDNA analysed by Bayesian inference and maximum parsimony. The resulting trees were well resolved at most nodes and demonstrated that the Microphalloidea, as represented by the present data-set, consists of three main clades corresponding to the families Lecithodendriidae, Microphallidae and Pleurogenidae + Prosthogonimidae. Interrelationships of taxa within each clade are considered; as a result of analysis of molecular and morphological data, Floridatrema Kinsella & Deblock, 1994 is synonymised with Maritrema Nicoll, 1907, Candidotrema Dollfus, 1951 with Pleurogenes Looss, 1896, and Schistogonimus Lühe, 1909 with Prosthogonimus Lühe, 1899. The taxonomic value of some morphological features, used traditionally for the differentiation of genera within the Lecithodendriidae and Prosthogonimidae, is reconsidered. Previous systematic schemes are discussed from the viewpoint of present results, and perspectives of future studies are outlined.

Animals↗

Combined psychophysiological assessment of ADHD: a pilot study of Bayesian probability approach illustrated by appraisal of ADHD in female college students.

Manifestations of ADHD are observed at both psychological and physiological levels and assessed via various psychometric, EEG, and imaging tests. However, no test is 100% accurate in its assessment of ADHD. This study introduces a stochastic assessment combining psychometric tests with previously reported (Consistency Index) and newly developed (Alpha Blockade Index) EEG-based physiological markers of ADHD. The assessment utilizes classical Bayesian inference to refine after each step the probability of ADHD of each individual. In a pilot study involving six college females with ADHD and six matched controls, the assessment achieved correct classification for all ADHD and non-ADHD participants. In comparison, the classification of ADHD versus non-ADHD participants was < 85% for any one of the tests separately. The procedure significantly improved the score separation between ADHD versus non-ADHD groups. The final average probabilities for ADHD were 76% for the ADHD group and 8% for the control group. These probabilities correlated (r = .87) with the Brown ADD scale and (r = .84) with the ADHD-Symptom Inventory used for the screening of the participants. We conclude that, although each separate test was not completely accurate, a combination of several tests classified correctly all ADHD and all non-ADHD participants. The application of the proposed assessment is not limited to the specific tests used in this study--the assessment represents a general paradigm capable of accommodating a variety of ADHD tests into a single diagnostic assessment.

Adult↗

BACUS: A Bayesian protocol for the identification of protein NOESY spectra via unassigned spin systems.

NMR frequency assignments are usually considered a prerequisite for the analysis of NOESY spectra, in turn required for the calculation of biomolecular structures. In contrast, as we propose here, relatively high numbers of unambiguous NOE identities can be consistently achieved in an automated manner by relying only on grouping resonances into connected spin systems. To achieve this goal, we have developed for proteins two protocols, SPI and BACUS, based on Bayesian inference. SPI (Grishaev and Llinás, 2002c) produces a list of the (1)H resonance frequencies from homo- and hetero-nuclear multidimensional spectra, grouped into effective spin systems. BACUS automatically establishes probabilistic identities of NOESY cross-peaks in terms of the chemical shifts provided by SPI. BACUS requires neither assignment of resonances nor an initial structural model. It successfully copes with chemical shift overlap and does so without cycling through 3D structure calculations. The method exploits the self-consistency of the NOESY graph by taking advantage of a network of J- as well as NOE-connected "reporter" protons sorted via SPI. BACUS was validated by tests on experimental NOESY data recorded for the col 2 and kringle 2 domains.

Bayes Theorem↗

A general approach to single-nucleotide polymorphism discovery.

Single-nucleotide polymorphisms (SNPs) are the most abundant form of human genetic variation and a resource for mapping complex genetic traits. The large volume of data produced by high-throughput sequencing projects is a rich and largely untapped source of SNPs (refs 2, 3, 4, 5). We present here a unified approach to the discovery of variations in genetic sequence data of arbitrary DNA sources. We propose to use the rapidly emerging genomic sequence as a template on which to layer often unmapped, fragmentary sequence data and to use base quality values to discern true allelic variations from sequencing errors. By taking advantage of the genomic sequence we are able to use simpler yet more accurate methods for sequence organization: fragment clustering, paralogue identification and multiple alignment. We analyse these sequences with a novel, Bayesian inference engine, POLYBAYES, to calculate the probability that a given site is polymorphic. Rigorous treatment of base quality permits completely automated evaluation of the full length of all sequences, without limitations on alignment depth. We demonstrate this approach by accurate SNP predictions in human ESTs aligned to finished and working-draft quality genomic sequences, a data set representative of the typical challenges of sequence-based SNP discovery.

Algorithms↗

Genome-wide analysis of single-nucleotide polymorphisms in human expressed sequences.

Single-nucleotide polymorphisms (SNPs) have been explored as a high-resolution marker set for accelerating the mapping of disease genes. Here we report 48,196 candidate SNPs detected by statistical analysis of human expressed sequence tags (ESTs), associated primarily with coding regions of genes. We used Bayesian inference to weigh evidence for true polymorphism versus sequencing error, misalignment or ambiguity, misclustering or chimaeric EST sequences, assessing data such as raw chromatogram height, sharpness, overlap and spacing, sequencing error rates, context-sensitivity and cDNA library origin. Three separate validations-comparison with 54 genes screened for SNPs independently, verification of HLA-A polymorphisms and restriction fragment length polymorphism (RFLP) testing-verified 70%, 89% and 71% of our predicted SNPs, respectively. Our method detects tenfold more true HLA-A SNPs than previous analyses of the EST data. We found SNPs in a large fraction of known disease genes, including some disease-causing mutations (for example, the HbS sickle-cell mutation). Our comprehensive analysis of human coding region polymorphism provides a public resource for mapping of disease genes (available at http://www.bioinformatics.ucla.edu/snp).

Base Sequence↗

Biased biological functions of horizontally transferred genes in prokaryotic genomes.

Horizontal gene transfer is one of the main mechanisms contributing to microbial genome diversification. To clarify the overall picture of interspecific gene flow among prokaryotes, we developed a new method for detecting horizontally transferred genes and their possible donors by Bayesian inference with training models for nucleotide composition. Our method gives the average posterior probability (horizontal transfer index) for each gene sequence, with a low horizontal transfer index indicating recent horizontal transfer. We found that 14% of open reading frames in 116 prokaryotic complete genomes were subjected to recent horizontal transfer. Based on this data set, we quantitatively determined that the biological functions of horizontally transferred genes, except mobile element genes, are biased to three categories: cell surface, DNA binding and pathogenicity-related functions. Thus, the transferability of genes seems to depend heavily on their functions.

Algorithms↗

Reconstructing population exposures from dose biomarkers: inhalation of trichloroethylene (TCE) as a case study.

Physiologically based pharmacokinetic (PBPK) modeling is a well-established toxicological tool designed to relate exposure to a target tissue dose. The emergence of federal and state programs for environmental health tracking and the availability of exposure monitoring through biomarkers creates the opportunity to apply PBPK models to estimate exposures to environmental contaminants from urine, blood, and tissue samples. However, reconstructing exposures for large populations is complicated by often having too few biomarker samples, large uncertainties about exposures, and large interindividual variability. In this paper, we use an illustrative case study to identify some of these difficulties, and for a process for confronting them by reconstructing population-scale exposures using Bayesian inference. The application consists of interpreting biomarker data from eight adult males with controlled exposures to trichloroethylene (TCE) as if the biomarkers were random samples from a large population with unknown exposure conditions. The TCE concentrations in blood from the individuals fell into two distinctly different groups even though the individuals were simultaneously in a single exposure chamber. We successfully reconstructed the exposure scenarios for both subgroups - although the reconstruction of one subgroup is different than what is believed to be the true experimental conditions. We were however unable to predict with high certainty the concentration of TCE in air.

Adult↗

Phylogeny and evolution of the major intrinsic protein family.

BACKGROUND INFORMATION: MIPs (major intrinsic proteins) form channels across biological membranes that control recruitment of water and small solutes such as glycerol and urea in all living organisms. Because of their widespread occurrence and large number, MIPs are a sound model system to understand evolutionary mechanisms underlying the generation of protein structural and functional diversity. With the recent increase in genomic projects, there is a considerable increase in the quantity and taxonomic range of MIPs in molecular databases. RESULTS: In the present study, I compiled more than 450 non-redundant amino acid sequences of MIPs from NCBI databases. Phylogenetic analyses using Bayesian inference reconstructed a statistically robust tree that allowed the classification of members of the family into two main evolutionary groups, the GLPs (glycerol-uptake facilitators or aquaglyceroporins) and the water transport channels or AQPs (aquaporins). Separate phylogenetic analyses of each of the MIP subfamilies were performed to determine the main groups of orthology. In addition, comparative sequence analyses were conducted to identify conserved signatures in the MIP molecule. CONCLUSIONS: The earliest and major gene duplication event in the history of the MIP family led to its main functional split into GLPs and AQPs. GLPs show typically one single copy in microbes (eubacteria, archaea and fungi), up to four paralogues in vertebrates and they are absent from plants. AQPs are usually single in microbes and show their greatest numbers and diversity in angiosperms and vertebrates. Functional recruitment of NOD26-like intrinsic proteins to glycerol transport due to the absence of GLPs in plants was highly supported. Acquisition of other MIP functions such as permeability to ammonia, arsenite or CO2 is restricted to particular MIP paralogues. Up to eight fairly conserved boxes were inferred in the primary sequence of the MIP molecule. All of them mapped on to one side of the channel except the conserved glycine residues from helices 2 and 5 that were found in the opposite side.

Amino Acid Sequence↗

CisModule: de novo discovery of cis-regulatory modules by hierarchical mixture modeling.

The regulatory information for a eukaryotic gene is encoded in cis-regulatory modules. The binding sites for a set of interacting transcription factors have the tendency to colocalize to the same modules. Current de novo motif discovery methods do not take advantage of this knowledge. We propose a hierarchical mixture approach to model the cis-regulatory module structure. Based on the model, a new de novo motif-module discovery algorithm, CisModule, is developed for the Bayesian inference of module locations and within-module motif sites. Dynamic programming-like recursions are developed to reduce the computational complexity from exponential to linear in sequence length. By using both simulated and real data sets, we demonstrate that CisModule is not only accurate in predicting modules but also more sensitive in detecting motif patterns and binding sites than standard motif discovery methods are.

Algorithms↗

Bayesian coclustering of Anopheles gene expression time series: study of immune defense response to multiple experimental challenges.

We present a method for Bayesian model-based hierarchical coclustering of gene expression data and use it to study the temporal transcription responses of an Anopheles gambiae cell line upon challenge with multiple microbial elicitors. The method fits statistical regression models to the gene expression time series for each experiment and performs coclustering on the genes by optimizing a joint probability model, characterizing gene coregulation between multiple experiments. We compute the model using a two-stage Expectation-Maximization-type algorithm, first fixing the cross-experiment covariance structure and using efficient Bayesian hierarchical clustering to obtain a locally optimal clustering of the gene expression profiles and then, conditional on that clustering, carrying out Bayesian inference on the cross-experiment covariance using Markov chain Monte Carlo simulation to obtain an expectation. For the problem of model choice, we use a cross-validatory approach to decide between individual experiment modeling and varying levels of coclustering. Our method successfully generates tightly coregulated clusters of genes that are implicated in related processes and therefore can be used for analysis of global transcript responses to various stimuli and prediction of gene functions.

Algorithms↗