Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Plastome evolution and phylogenomic relationships in Ajuga (Lamiaceae, Ajugoideae).

BACKGROUND: Ajuga is currently known to include approximately 69 species, with a combined distribution extending throughout Eurasia, Africa, and Australia. Its popularity and significance are largely based on an extensive history of medicinal and horticultural use. It is divided into two sections based on morphological characters, and this sectional classification is also reflected in pronounced geographic patterns. Although previous studies have largely focused on Ajuga sect. Ajuga in East Asia, A. sect. Chamaepithys, which ranges from the Mediterranean to Central Asia, remains insufficiently sampled, thereby limiting a comprehensive understanding of infrageneric sectional relationships within the genus. Here, we generated complete plastid genomes for 12 species representing both sections of the genus and used these data to characterize plastome structure and infer evolutionary relationships. RESULTS: In this study, 21 Ajuga plastomes were analyzed, including 12 newly sequenced plastomes and 9 previously published plastomes representing 19 species. Comparative analyses showed that all plastomes exhibited a highly conserved quadripartite structure, with genome sizes ranging from 149,963 to 150,740 bp and GC contents varying from 38.2% to 38.3%. Each plastome contained 133 genes, including 88 protein-coding genes, 37 transfer RNA genes, and 8 ribosomal RNA genes. The boundaries between the inverted repeat (IR) and single-copy (SC) regions were also highly conserved across species. In addition, 796 simple sequence repeats (SSRs), 874 long repeat sequences (LRSs), and 12 highly variable regions (ccsA-ndhD, ndhF-rpl32, petA-psbJ, rpl32-trnL-UAG, rps2-rpoC2, trnH-GUG-psbA, trnK-UUU-rps16, trnP-UGG-psaJ, trnT-UGU-trnL-UAA, ycf15-trnL-CAA, ndhF, and ycf1) were identified among the 21 plastomes. Phylogenetic analyses based on four datasets and conducted using Maximum Likelihood and Bayesian Inference recovered two major clades corresponding to the traditionally recognized sectional classification, with one distributed from the Mediterranean to Central Asia and the other in East Asia. CONCLUSION: This study represents the most comprehensive plastome-based sampling of Ajuga to date, including representative species from the Mediterranean, Central Asia, and East Asia. Our results have significantly enhanced our understanding of its infrageneric relationships. The plastome resources generated in this study provide a valuable foundation for future research on species delimitation, phylogeny, and the evolutionary history of Ajuga.

Phylogeny↗

Evolutionary HMMs: a Bayesian approach to multiple alignment.

MOTIVATION: We review proposed syntheses of probabilistic sequence alignment, profiling and phylogeny. We develop a multiple alignment algorithm for Bayesian inference in the links model proposed by Thorne et al. (1991, J. Mol. Evol., 33, 114-124). The algorithm, described in detail in Section 3, samples from and/or maximizes the posterior distribution over multiple alignments for any number of DNA or protein sequences, conditioned on a phylogenetic tree. The individual sampling and maximization steps of the algorithm require no more computational resources than pairwise alignment. METHODS: We present a software implementation (Handel) of our algorithm and report test results on (i) simulated data sets and (ii) the structurally informed protein alignments of BAliBASE (Thompson et al., 1999, Nucleic Acids Res., 27, 2682-2690). RESULTS: We find that the mean sum-of-pairs score (a measure of residue-pair correspondence) for the BAliBASE alignments is only 13% lower for Handelthan for CLUSTALW(Thompson et al., 1994, Nucleic Acids Res., 22, 4673-4680), despite the relative simplicity of the links model (CLUSTALW uses affine gap scores and increased penalties for indels in hydrophobic regions). With reference to these benchmarks, we discuss potential improvements to the links model and implications for Bayesian multiple alignment and phylogenetic profiling. AVAILABILITY: The source code to Handelis freely distributed on the Internet at http://www.biowiki.org/Handel under the terms of the GNU Public License (GPL, 2000, http://www.fsf.org./copyleft/gpl.html).

Algorithms↗

Bayesian analysis for a single 2 x 2 table.

The simple comparison of two binomial populations is frequently of interest in epidemiology when the domains are large. For small domains, however, there are no exact methods except Fisher's exact test. A basic problem, therefore, is to compare two populations by assessing the difference between the proportions of individuals who possess a characteristic in the first and second populations. When there is prior information, we take the proportions to have independent conjugate beta distributions with known parameters, thereby facilitating a Bayesian analysis. We consider Bayesian inference on functions of the proportions, and the three most common scalar measures used in epidemiology and health services research, namely relative risk, odds ratio and attributable risk. We develop the highest density regions (both exact and approximate) for relative risk, odds ratio and attributable risk. In addition, we consider the Bayes factor for testing whether the model with a common proportion holds rather than one with distinct proportions. Using data from the population-based Worcester Heart Attack Study, we apply our methodology to study gender differences in the therapeutic management of patients with acute myocardial infarction (AMI) by selected demographic and clinical characteristics. The Bayes factor, the approximate and exact intervals generally suggest that there are no substantial differences in the pharmacologic management of males and females hospitalized with AMI.

Adult↗

Identifying the types of missingness in quality of life data from clinical trials.

This paper discusses methods of identifying the types of missingness in quality of life (QOL) data in cancer clinical trials. The first approach involves collecting information on why the QOL questionnaires were not completed. Based on the reasons provided one may be able to distinguish the mechanisms causing missing data. The second approach is to model the missing data mechanism and perform hypothesis testing to determine the missing data processes. Two methods of testing if missing data are missing completely at random (MCAR) are presented and applied to incomplete longitudinal QOL data obtained from international multi-centre cancer clinical trials. The first method (Ridout, 1991) is based on a logistic regression and the second method (Park and Davis, 1993) is based on an adaptation of weighted least squares. In one application (advanced breast cancer) missing data was not likely to be MCAR. In the second application (adjuvant breast cancer) the missing mechanism was dependent on the QOL scale under study. MCAR and missing at random (MAR) have distinct consequences for data analysis. Therefore it is relevant to distinguish between them. However, if either MCAR or MAR hold, likelihood or Bayesian inferences can be based solely on the observed data, although for MAR, depending on the research question, modelling the dropout mechanism may still be necessary. Distinguishing between MAR and missing not at random (MNAR) is not trivial and relies on fundamentally untestable assumptions.

Clinical Trials as Topic↗

Assessing heterogeneity and correlation of paired failure times with the bivariate frailty model.

We consider bivariate survival times for heterogeneous populations, where heterogeneity induces deviations in an individual's risk of an event as well as associations between survival times. The heterogeneity is characterized by a bivariate frailty model. We measure the heterogeneity effects through deviations associated with hazard functions and an association function defined through the conditional hazard functions: the cross-ratio function proposed by Oakes. We show how the deviation and association measures are determined by the frailty distribution. A Gibbs sampling method is developed for Bayesian inferences on regression coefficients, frailty parameters and the heterogeneity measures. The method is applied to a mental health care data set.

Algorithms↗

A bayesian analysis for spatial processes with application to disease mapping.

In epidemiology, maps of disease rates and disease risk provide a spatial perspective for researching disease aetiology. For rare diseases or when the population base is small, the rate and risk estimates may be unstable. We propose using a Bayesian analysis based on the conditional autoregressive (CAR) process that will spatially smooth disease rates or risk estimates by allowing each site to 'borrow strength' from its neighbours. Covariates may be included in the model in such a way as to establish a possible association between risk factors and disease incidence. Bayesian inferences are implemented from a direct resampling scheme where large samples are generated from the various posterior distributions. The methodology is demonstrated with a simulation that assesses the effect of sample size and the model parameters on inferences for the parameters. Our approach is also used to spatially smooth district lip cancer rates in Scotland using the CAR model with a covariate that allows for exposure to sunlight.

Bayes Theorem↗

Genetic variance components analysis for binary phenotypes using generalized linear mixed models (GLMMs) and Gibbs sampling.

The common complex diseases such as asthma are an important focus of genetic research, and studies based on large numbers of simple pedigrees ascertained from population-based sampling frames are becoming commonplace. Many of the genetic and environmental factors causing these diseases are unknown and there is often a strong residual covariance between relatives even after all known determinants are taken into account. This must be modelled correctly whether scientific interest is focused on fixed effects, as in an association analysis, or on the covariances themselves. Analysis is straightforward for multivariate Normal phenotypes, but difficulties arise with other types of trait. Generalized linear mixed models (GLMMs) offer a potentially unifying approach to analysis for many classes of phenotype including multivariate Normal traits, binary traits, and censored survival times. Markov Chain Monte Carlo methods, including Gibbs sampling, provide a convenient framework within which such models may be fitted. In this paper, Bayesian inference Using Gibbs Sampling (a generic Gibbs sampler; BUGS) is used to fit GLMMs for multivariate Normal and binary phenotypes in nuclear families. BUGS is easy to use and readily available. We motivate a suitable model structure for Normal phenotypes and show how the model extends to binary traits. We discuss parameter interpretation and statistical inference and show how to circumvent a number of important theoretical and practical problems that we encountered. Using simulated data we show that model parameters seem consistent and appear unbiased in smaller data sets. We illustrate our methods using data from an ongoing cohort study.

Binomial Distribution↗

Modelling the cumulative risk for a false-positive under repeated screening events.

Screening examinations are widely utilized in detecting the presence of medical disorders, for instance, screening mammograms and clinical breast examinations for detection of breast cancer. Such procedures are invaluable in enabling early treatment but produce the possibilities of false-positive and false-negative diagnoses. Focusing on false-positive results, with increasing number of screening events, it is clear that the risk of a false-positive increases. The objective of this paper is to quantify the cumulative risk associated with repeated screening. We provide a very general framework within which to investigate this risk, both at the population and the individual level. The latter allows incorporation of evolving patient medical history to permit individualized assessment of risk. We model cumulative risk in terms of the number of screening events until first false-positive. We develop models which are essentially familiar actuarial models for life table data adding a Cox regression to enable individual level modelling. Because it offers several advantages, we employ a Bayesian inference framework and apply our modelling to the analysis of 9773 screening mammograms collected from 2227 women at an HMO serving nearly 300000 adults in and around Boston, MA.

Adult↗

Variance components analysis for pedigree-based censored survival data using generalized linear mixed models (GLMMs) and Gibbs sampling in BUGS.

Complex human diseases are an increasingly important focus of genetic research. Many of the determinants of these diseases are unknown and there is often a strong residual covariance between relatives even when all known genetic and environmental factors have been taken into account. This must be modeled correctly whether scientific interest is focused on fixed effects, as in an association analysis, or on the covariance structure itself. Analysis is straightforward for multivariate normally distributed traits, but difficulties arise with other types of trait. Generalized linear mixed models (GLMMs) offer a potentially unifying approach to analysis for many classes of phenotype including right censored survival times. This includes age-at-onset and age-at-death data and a variety of other censored traits. Markov chain Monte Carlo (MCMC) methods, including Gibbs sampling, provide a convenient framework within which such GLMMs may be fitted. In this paper, we use BUGS ("Bayesian inference using Gibbs sampling": a readily available, generic Gibbs sampler) to fit GLMMs for right-censored survival times in nuclear and extended families. We discuss parameter interpretation and statistical inference, and show how to circumvent a number of important theoretical and practical problems. Using simulated data, we show that model parameters are consistent. We further illustrate our methods using data from an ongoing cohort study. Finally, we propose that the random effects associated with a genetic component of variance (e.g., sigma(2)(A)) in a GLMM may be regarded as an adjusted "phenotype" and used as input to a conventional model-based or model-free linkage analysis. This provides a simple way to conduct a linkage analysis for a trait reflected in a right-censored survival time while comprehensively adjusting for observed confounders at the level of the individual and latent environmental effects shared across families.

Bayes Theorem↗

Insights Into the Structural Features, Codon Usage Patterns, and Phylogenetic Analysis in Neoniphon argenteus (Teleostei: Holocentriformes) Based on Complete Mitochondrial Genome.

Neoniphon argenteus, a widely distributed nocturnal coral reef fish in the family Holocentridae, plays an important role in maintaining coral reef ecosystem health, yet its phylogenetic position remains poorly resolved. To bridge this gap, we sequenced and analyzed the complete mitochondrial genome of a specimen from the South China Sea to characterize its structural features, codon usage patterns, and phylogenetic relationships. The 16,569 bp mitogenome (GenBank: PP190474.1) encodes 13 protein-coding genes (PCGs), 22 tRNAs, two rRNAs, and two non-coding regions, exhibiting a distinct A + T bias. All tRNAs fold into typical cloverleaf secondary structures except tRNA-Ser (AGN), which lacks the dihydrouridine (DHU) arm. The control region contains palindromic motifs (TACAT/ATGTA) capable of forming hairpin structures and five conserved sequence blocks, whereas the OL region harbors a conserved 5'-GCCGG-3' motif. RSCU analysis revealed 31 frequently used codons (RSCU > 1) with a pronounced preference for A/C-ending codons. The ΔRSCU method identified 10 candidate optimal codons (GCA, CAA, GAA, GGA, AUU, CUA, CCA, CGA, ACA, and GUC). Selection pressure analysis using EasyCodeML and site-specific models indicated that all PCGs are predominantly under purifying selection, with no significant evidence of pervasive positive selection. ND6 exhibited elevated pairwise Ka/Ks ratios (mean = 1.209 ± 0.047), consistent with reduced selective constraint rather than adaptive evolution. Phylogenetic analysis of 19 Holocentriformes species using maximum likelihood and Bayesian inference with partitioned models based on 13 PCGs and two rRNA genes (12S and 16S) assigned all taxa to two well-supported subfamilies (Holocentrinae and Myripristinae). Within Holocentrinae, Neoniphon species form a monophyletic clade nested within a paraphyletic Sargocentron, suggesting that the genus Sargocentron as currently defined is not monophyletic. This study provides useful baseline molecular data for further exploration of the evolutionary history of N. argenteus and other members of Holocentriformes.

Holocentridae↗

Bayesian technique for investigating linearity in event-related BOLD fMRI.

Event-related BOLD fMRI data is modeled as a linear time-invariant system. Together with Bayesian inference techniques, a statistical test is developed for rigorously detecting linearity/nonlinearity in the BOLD response system. The test is applied to data collected from eight subjects using an event-related paradigm with a switching checkerboard as the visual stimulus. Analyzed as a group, the results clearly find the response to be nonlinear. When each subject is analyzed individually, however, the results are predominantly nonlinear, but there is some evidence to suggest that there may be a crossover from a linear to a nonlinear regime and vice versa. This could be important when estimating physiological parameters for individuals. Additionally, estimates of the hemodynamic response function and corresponding response were obtained, but there was no consistent appearance of a poststimulus undershoot in the event-related BOLD response.

Adult↗

Comparison of the information in two lung function experiments.

The amount of ventilation relative to perfusion (the ventilation-perfusion ratio) received by the lung is a useful indicator of the efficiency of lung function. Two alternative techniques for recovering the ventilation-perfusion ratio are outlined. While both techniques rely on the use of inert gases, one is well established and the other is only in a developmental stage. This paper focuses on a comparison of the amount of statistical information provided by these two techniques about the ventilation-perfusion ratio. The criterion applied here for measuring amount of information has roots in communication theory and uses ideas inherent to Bayesian inference.

Bayes Theorem↗

The pattern of variation in centipede segment number as an example of developmental constraint in evolution

The range of animal morphologies observed in nature is partly determined by natural selection. However, there is no agreement yet regarding whether it is also partly determined by developmental constraint. Testing for the effects of constraint has been difficult due to the lack of both an appropriate null model and a sufficiently simple system capable of yielding unambiguous results regarding the model's plausibility. Here we examine the case of variation in segment number in geophilomorph centipedes. Curiously, while this ranges between 29 and 191, there are no species in which an even number of segments is observed, in contrast to about 1000 species with odd numbers of segments. It seems unlikely that this distribution of character values is determined by selection alone. Using an approach based on Bayesian inference, we attempt to quantify the probability of obtaining the observed distribution of values given a null model in which developmental constraint is absent. Since this probability is in the region of 10(-20), we conclude that constraint must be involved. We discuss various implications of this conclusion, and comment on the unexpected absence of neoteny and progenesis in centipede evolution. Copyright 1999 Academic Press.

Journal Article↗

On the probability model for asthma attacks.

In environmental epidemiology, the impact of environmental agents on symptoms or health status is of interest. This influence is described quantitatively in the theory of Whittemore & Keller (1979). They formulated a logistic model for individuals that is useful in evaluation of panel studies in which each participant protocols whether he does or does not have a certain symptom each day. In the present paper an equation for the prevalence of symptoms in the study population that is defined as the fraction of symptomatic subjects is deduced from the model for individuals. The model for the aggregated quantity depends on the individuals' parameters in a nonlinear manner. The relationship between the individual-based model and the corresponding population-based model is illustrated by means of a simulated panel. Bayesian estimates of the parameters are calculated and compared for both approaches. Bayesian inference enables to apply the prevalence model to a population of non-identical individuals. For such a heterogeneous population, we observe an attenuation of environmental effects on the aggregated symptom prevalence in comparison to the individual-based approach. The presented theory is applicable not only to panel studies but also in time-series analysis of prevalences and incidences.

Asthma↗

Bayesian estimation of dynamical systems: an application to fMRI.

This paper presents a method for estimating the conditional or posterior distribution of the parameters of deterministic dynamical systems. The procedure conforms to an EM implementation of a Gauss-Newton search for the maximum of the conditional or posterior density. The inclusion of priors in the estimation procedure ensures robust and rapid convergence and the resulting conditional densities enable Bayesian inference about the model parameters. The method is demonstrated using an input-state-output model of the hemodynamic coupling between experimentally designed causes or factors in fMRI studies and the ensuing BOLD response. This example represents a generalization of current fMRI analysis models that accommodates nonlinearities and in which the parameters have an explicit physical interpretation. Second, the approach extends classical inference, based on the likelihood of the data given a null hypothesis about the parameters, to more plausible inferences about the parameters of the model given the data. This inference provides for confidence intervals based on the conditional density.

Bayes Theorem↗

A Bayesian approach to Weibull survival models--application to a cancer clinical trial.

In this paper we outline a class of fully parametric proportional hazards models, in which the baseline hazard is assumed to be a power transform of the time scale, corresponding to assuming that survival times follow a Weibull distribution. Such a class of models allows for the possibility of time varying hazard rates, but assumes a constant hazard ratio. We outline how Bayesian inference proceeds for such a class of models using asymptotic approximations which require only the ability to maximize the joint log posterior density. We apply these models to a clinical trial to assess the efficacy of neutron therapy compared to conventional treatment for patients with tumours of the pelvic region. In this trial there was prior information about the log hazard ratio both in terms of elicited clinical beliefs and the results of previous studies. Finally, we consider a number of extensions to this class of models, in particular the use of alternative baseline functions, and the extension to multi-state data.

Bayes Theorem↗

Probability and the patient state space.

This paper describes work to develop a model-based system to support clinical decision-making. In previous articles, we have developed (from 695 measurement sets obtained from 148 patients) a physiologic state classification based on a set of 11 cardiovascular and metabolic measurements. There is an R or reference state, for stable ICU patients. Patients under (operative, traumatic, or compensated septic) stress, or with (septic or hepatic) metabolic, respiratory, or cardiac insufficiency are in the A, B, C, or D states, respectively. We wished to make the state easier to measure and eventually available continuously, automatically, and noninvasively, as well as reflecting a wider group of bodily systems. The 5 centers define a 4 dimensional affine subspace, designated the cardiovascular state space. Using eigenvector analysis, we have found four new derived physiologic variables CV1, CV2, CV3, and CV4 that span the state space. We have fit sets of linear regression equations that allow the patient's position in the state space, and therefore his state, to be determined from more easily obtainable sets of measurements. Further, we selected 1966 measurement sets from 512 patients at two hospitals. We used the data from 250 of these patients to define 13 prototypical types, namely survivors and deaths from various combinations of sepsis, cardiogenic decompensation, cirrhosis, and pneumonitis, following trauma or general surgery. For any future patient, the statistical theory of Bayesian inference allows one to infer back from the measurements observed to the probability of his being of any of these types and of surviving or dying. We used this method to predict the outcome of the other 262 patients, prospectively. Statistically, the predictions of survival or death were not significantly different from the actual. For individual patients, the method predicts a clinical course that closely follows the actual episodes in their history. These results confirm and explain the validity of the concept of the patient state and make the state easier to compute. The patient state and the probability plot together help to stage, select, and evaluate therapy. They do not replace the clinician's judgement, but rather are tools that help the clinician to exercise judgement.

Adult↗

Medical expert systems based on causal probabilistic networks.

Causal probabilistic networks (CPNs) offer new methods by which you can build medical expert systems that can handle all types of medical reasoning within a uniform conceptual framework. Based on the experience from a commercially available system and a couple of large prototype systems, it appears that CPNs are now an attractive alternative to other methods. A CPN is an intensional model of a domain, and it is therefore conceptually much closer to qualitative reasoning systems and to simulation systems than to rule-based or logic-based systems. Recent progress in Bayesian inference in networks has yielded computationally efficient methods. The inference method used follows the fundamental axioms of probability theory, and gives a sound framework for causal and diagnostic (deductive and abductive) reasoning under uncertainty. Experience with the prototypes indicates that it may be possible to use decision theory as a rational approach to test planning and therapy planning. The way in which knowledge is acquired and represented in CPNs makes it easy to express 'deep knowledge' for example in the form of physiological models, and the facilities for learning make it possible to make a smooth transition from expert opinion to statistics based on empirical data.

Artificial Intelligence↗