Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Predictive inference, causal reasoning, and model assessment in nonparametric Bayesian analysis: a case study.

This paper continues our earlier analysis of a data set on acute ear infections in small children, presented in Andreev and Arjas (1998). The main goal here is to provide a method, based on the use of predictive distributions, for assessing the possible causal influence which the type of day care will have on the incidence of ear infections. A closely related technique is used for the assessment of the nonparametric Bayesian intensity model applied in the paper. Two graphical methods, supported by formal tests, are suggested for this purpose.

Algorithms↗

Bayesian detection and modeling of spatial disease clustering.

Many current statistical methods for disease clustering studies are based on a hypothesis testing paradigm. These methods typically do not produce useful estimates of disease rates or cluster risks. In this paper, we develop a Bayesian procedure for drawing inferences about specific models for spatial clustering. The proposed methodology incorporates ideas from image analysis, from Bayesian model averaging, and from model selection. With our approach, we obtain estimates for disease rates and allow for greater flexibility in both the type of clusters and the number of clusters that may be considered. We illustrate the proposed procedure through simulation studies and an analysis of the well-known New York leukemia data.

Bayes Theorem↗

Clinical inferences and decisions--III. Utility assessment and the Bayesian decision model.

It is accepted that errors of misclassifications, however small, can occur in clinical decisions but it cannot be assumed that the importance associated with false positive errors is the same as that for false negatives. The relative importance of these two types of error is frequently implied by a decision maker in the different weighting factors or utilities he assigns to the alternative consequences of his decisions. Formal procedures are available by which it is possible to make explicit in numerical form the value or worth of the outcome of a decision process. The two principal methods are described for generating utilities as associated with clinical decisions. The concept and application of utility is then expanded from a unidimensional to a multidimensional problem where, for example, one variable may be state of health and another monetary assets. When combined with the principles of subjective probability and test criterion selection outlined in Parts I and II of this series, the consequent use of utilities completes the framework upon which the general Bayesian model of clinical decision making is based. The five main stages in this general decision making model are described and applications of the model are illustrated with clinical examples from the field of ophthalmology. These include examples for unidimensional and multidimensional problems which are worked through in detail to illustrate both the principles and methodology involved in a rationalized normative model of clinical decision making behaviour.

Bayes Theorem↗

Multiple-trait Gibbs sampler for animal models: flexible programs for Bayesian and likelihood-based (co)variance component inference.

A set of FORTRAN programs to implement a multiple-trait Gibbs sampling algorithm for (co)variance component inference in animal models (MTGSAM) was developed. The MTGSAM programs are available to the public. The programs support models with correlated genetic effects and arbitrary numbers of covariates, fixed effects, and independent random effects for each trait. Any combination of missing traits is allowed. The programs were used to estimate variance components for 50 replicates of simulated data. Each replicate consisted of 50 animals of each sex in each of four generations, for 400 animals in each replicate for two traits. For MTGSAM, informative prior distributions for variance components were inverted Wishart random variables with 10 df and means equal to the simulation parameters. A total of 15,000 Gibbs sampling rounds were completed for each replicate, with 2,000 rounds discarded for burn-in. For multiple-trait derivative free restricted maximum likelihood (MTDFREML), starting values for the variance components were the simulation parameters. Averages of posterior mean of variance components estimated using MTGSAM with informative and flat prior distributions for variance components and REML estimates obtained using MTDFREML indicated that all three methods were empirically unbiased. Correlations between estimates from MTGSAM using flat priors and MTDFREML all exceeded .99.

Algorithms↗

ReGAIN: a bioinformatics platform for assessing probabilistic co-occurrence between resistance genes in bacterial pathogens.

MOTIVATION: Multidrug-resistant bacterial pathogens continue to rise globally, yet scalable methods are needed to infer how resistance determinants co-occur across pathogen populations and to quantify conditional dependencies underlying co-occurrence and shared genetic context. RESULTS: We present ReGAIN (Resistance Gene Association and Inference Network), an open-source platform that applies Bayesian network structure learning to infer probabilistic, conditional dependency relationships among antibiotic resistance, heavy metal tolerance, stress response, and virulence determinants in bacteria. In contrast to pairwise co-occurrence analyses, ReGAIN reports conditional probabilities, relative risks, and absolute risk differences with confidence intervals to prioritize candidate relationships for downstream prioritization. Applied across ESKAPEE pathogens, ReGAIN recapitulated established resistance gene relationships and identified additional candidate patterns consistent with co-selection and shared genetic context. Together, these results support scalable, reproducible population-wide analysis of resistance networks for surveillance, comparative genomics and epidemiology. AVAILABILITY: ReGAIN analyses are performed using Python v3.11.5 and R v4.4.1 and is available as open-source software through Bioconda at {https://anaconda.org/bioconda/regain-cli}. Source code and documentation can be found at {https://github.com/ERBringHorvath/regain_CLI}. All genomes used in this publication were downloaded from the National Center for Biotechnology Information database. Large supplementary tables and results data from the ESKAPEE pathogen example network analyses can be downloaded from https://figshare.com/articles/dataset/ReGAIN_command_line_software_and_supplemental_figures_/28959431.

Computational Biology↗

Exponential family models and statistical genetics.

This article describes the evolution of applied exponential family models, starting at 1972, the year of publication of the seminal papers on generalized linear models and on Cox regression, and leading to multivariate (i) marginal models and inference based on estimating equations and (ii) random effects models and Bayesian simulation-based posterior inference. By referring to recent work in genetic epidemiology, on semiparametric methods for linkage analysis and on transmission/disequilibrium tests for haplotype transmission this paper illustrates the potential for the recent advances in applied probability and statistics to contribute to new and unified tools for statistical genetics. Finally, it is emphasized that there is a need for well-defined postgraduate education paths in medical statistics in the year 2000 and thereafter.

Biometry↗

An alternative to null-hypothesis significance tests.

The statistic p(rep) estimates the probability of replicating an effect. It captures traditional publication criteria for signal-to-noise ratio, while avoiding parametric inference and the resulting Bayesian dilemma. In concert with effect size and replication intervals, p(rep) provides all of the information now used in evaluating research, while avoiding many of the pitfalls of traditional statistical inference.

Bayes Theorem↗

Bayesian detection of periodic mRNA time profiles without use of training examples.

BACKGROUND: Detection of periodically expressed genes from microarray data without use of known periodic and non-periodic training examples is an important problem, e.g. for identifying genes regulated by the cell-cycle in poorly characterised organisms. Commonly the investigator is only interested in genes expressed at a particular frequency that characterizes the process under study but this frequency is seldom exactly known. Previously proposed detector designs require access to labelled training examples and do not allow systematic incorporation of diffuse prior knowledge available about the period time. RESULTS: A learning-free Bayesian detector that does not rely on labelled training examples and allows incorporation of prior knowledge about the period time is introduced. It is shown to outperform two recently proposed alternative learning-free detectors on simulated data generated with models that are different from the one used for detector design. Results from applying the detector to mRNA expression time profiles from S. cerevisiae showsthat the genes detected as periodically expressed only contain a small fraction of the cell-cycle genes inferred from mutant phenotype. For example, when the probability of false alarm was equal to 7%, only 12% of the cell-cycle genes were detected. The genes detected as periodically expressed were found to have a statistically significant overrepresentation of known cell-cycle regulated sequence motifs. One known sequence motif and 18 putative motifs, previously not associated with periodic expression, were also over represented. CONCLUSION: In comparison with recently proposed alternative learning-free detectors for periodic gene expression, Bayesian inference allows systematic incorporation of diffuse a priori knowledge about, e.g. the period time. This results in relative performance improvements due to increased robustness against errors in the underlying assumptions. Results from applying the detector to mRNA expression time profiles from S. cerevisiae include several new findings that deserve further experimental studies.

Algorithms↗

Algorithms for Bayesian background-subtracted Fourier darkfield imaging.

Formal consideration of prior information on the Fourier amplitude of background contrast in an image, using the same Bayesian principles of statistical inference which underlie thermodynamics, allows one to subtract background without favoring only selected parts of frequency space. Without the bias in frequency space which causes periodicity bleeding and mars literal interpretation of Fourier-filtered images, the shape transform of aperiodic objects can be left intact. Algorithms for Bayesian background subtraction from one- and two-dimensional images are presented which further consider, in ad hoc fashion, one's uncertainty about background amplitude. The results help explain the reported success of Fourier truncation, and indicate that Bayesian background-subtracted images can minimize root-mean-square image error, as well as periodicity bleeding, in comparison to Fourier-filtered and Fourier-truncated alternatives.

Algorithms↗

Cross-species analysis of biological networks by Bayesian alignment.

Complex interactions between genes or proteins contribute a substantial part to phenotypic evolution. Here we develop an evolutionarily grounded method for the cross-species analysis of interaction networks by alignment, which maps bona fide functional relationships between genes in different organisms. Network alignment is based on a scoring function measuring mutual similarities between networks, taking into account their interaction patterns as well as sequence similarities between their nodes. High-scoring alignments and optimal alignment parameters are inferred by a systematic Bayesian analysis. We apply this method to analyze the evolution of coexpression networks between humans and mice. We find evidence for significant conservation of gene expression clusters and give network-based predictions of gene function. We discuss examples where cross-species functional relationships between genes do not concur with sequence similarity.

Algorithms↗

Phylogeny of the Colubroidea (Serpentes): new evidence from mitochondrial and nuclear genes.

The Colubroidea contains over 85% of all the extant species of snakes and is recognized as monophyletic based on morphological and molecular data. Using DNA sequences (cyt b, c-mos) from 100 species we inferred the phylogeny of colubroids with special reference to the largest family, the Colubridae. Tree inference was obtained using Bayesian, likelihood, and parsimony methods. All analyses produced five major groups, the Pareatidae, Viperidae, Homalopsidae, the Elapidae, and the Colubridae. The specific content of the latter two groups has been altered to accommodate evolutionary history and to yield a more stable taxonomy. We propose an updated classification based on the reallocation of species as indicated by our inferred phylogeny.

Animals↗

Comparing ARG Inference Methods Under Transmission of Reproductive Success: Tree Imbalance Matters.

Inferring coalescent trees from genomic data has become a major subject in population genetics, particularly with the recent advances in tree sequence reconstruction methods. However, it remains unclear how well these methods perform for imbalanced genealogies. Such imbalances can arise from processes such as cultural transmission of reproductive success (CTRS) or positive selection. Using simulated genomic data, we benchmarked three major software packages, SINGER, Relate, and tsinfer, by comparing the imbalance of reconstructed trees by these methods with that of the true simulated trees, for three indices that quantify this imbalance. The three methods performed well under scenarios yielding balanced trees. However, their accuracy declined as imbalance increased. Performances also varied with mutation rate, recombination rate, and sample size. This study opens possibilities for applying these methods to infer CTRS or positive selection in large-scale genomic datasets, using simulation-based inference such as approximate Bayesian computation.

Models, Genetic↗

Multivariable modeling of radiotherapy outcomes, including dose-volume and clinical factors.

PURPOSE: The probability of a specific radiotherapy outcome is typically a complex, unknown function of dosimetric and clinical factors. Current models are usually oversimplified. We describe alternative methods for building multivariable dose-response models. METHODS: Representative data sets of esophagitis and xerostomia are used. We use a logistic regression framework to approximate the treatment-response function. Bootstrap replications are performed to explore variable selection stability. To guard against under/overfitting, we compare several analytical and data-driven methods for model-order estimation. Spearman's coefficient is used to evaluate performance robustness. Novel graphical displays of variable cross correlations and bootstrap selection are demonstrated. RESULTS: Bootstrap variable selection techniques improve model building by reducing sample size effects and unveiling variable cross correlations. Inference by resampling and Bayesian approaches produced generally consistent guidance for model order estimation. The optimal esophagitis model consisted of 5 dosimetric/clinical variables. Although the xerostomia model could be improved by combining clinical and dose-volume factors, the improvement would be small. CONCLUSIONS: Prediction of treatment response can be improved by mixing clinical and dose-volume factors. Graphical tools can mitigate the inherent complexity of multivariable modeling. Bootstrap-based variable selection analysis increases the reliability of reported models. Statistical inference methods combined with Spearman's coefficient provide an efficient approach to estimating optimal model order.

Carcinoma, Non-Small-Cell Lung↗

Differentiating between hypotheses of lineage sorting and introgression in New Zealand alpine cicadas (Maoricicada Dugdale).

Lineage sorting and introgression can lead to incongruence among gene phylogenies, complicating the inference of species trees for large groups of taxa that have recently and rapidly radiated. In addition, it can be difficult to determine which of these processes is responsible for this incongruence. We explore these issues with the radiation of New Zealand alpine cicadas of the genus Maoricicada Dugdale. Gene trees were estimated from four putative independent loci: mitochondrial DNA (2274 nucleotides), elongation factor 1-alpha (1275 nucleotides), period (1709 nucleotides), and calmodulin (678 nucleotides). We reconstructed phylogenies using maximum likelihood and Bayesian methods from 44 individuals representing the 19 species and subspecies of Maoricicada and two outgroups. Species-level relationships were reconstructed using a novel extension of gene tree parsimony, whereby gene trees were weighted by their Bayesian posterior probabilities. The inferred gene trees show marked incongruence in the placement of some taxa, especially the enigmatic forest and scrub dwelling species, M. iolanthe. Using the species tree estimated by gene tree parsimony, we simulated coalescent gene trees in order to test the null hypothesis that the nonrandom placement of M. iolanthe among gene trees has arisen by chance. Under the assumptions of constant population size, known generation time, and panmixia, we were able to reject this null hypothesis. Furthermore, because the two alternative placements of M. iolanthe are in each case with species that share a similar song structure, we conclude that it is more likely that an ancient introgression event rather than lineage sorting has caused this incongruence.

Animals↗

Phylogenetic MCMC algorithms are misleading on mixtures of trees.

Markov chain Monte Carlo (MCMC) algorithms play a critical role in the Bayesian approach to phylogenetic inference. We present a theoretical analysis of the rate of convergence of many of the widely used Markov chains. For N characters generated from a uniform mixture of two trees, we prove that the Markov chains take an exponentially long (in N) number of iterations to converge to the posterior distribution. Nevertheless, the likelihood plots for sample runs of the Markov chains deceivingly suggest that the chains converge rapidly to a unique tree. Our results rely on novel mathematical understanding of the log-likelihood function on the space of phylogenetic trees. The practical implications of our work are that Bayesian MCMC methods can be misleading when the data are generated from a mixture of trees. Thus, in cases of data containing potentially conflicting phylogenetic signals, phylogenetic reconstruction should be performed separately on each signal.

Algorithms↗

A population-based Bayesian approach to the minimal model of glucose and insulin homeostasis.

The minimal model was proposed in the late 1970s by Bergman et al. (Am. J. Physiol. 1979; 236(6):E667) as a powerful model consisting of three differential equations describing the glucose and insulin kinetics of a single individual. Considering the glucose and insulin simultaneously, the minimal model is a highly ill-posed estimation problem, where the reconstruction most often has been done by non-linear least squares techniques separately for each entity. The minimal model was originally specified for a single individual and does not combine several individuals with the advantage of estimating the metabolic portrait for a whole population. Traditionally it has been analysed in a deterministic set-up with only error terms on the measurements. In this work we adopt a Bayesian graphical model to describe the coupled minimal model that accounts for both measurement and process variability, and the model is extended to a population-based model. The estimation of the parameters are efficiently implemented in a Bayesian approach where posterior inference is made through the use of Markov chain Monte Carlo techniques. Hereby we obtain a powerful and flexible modelling framework for regularizing the ill-posed estimation problem often inherited in coupled stochastic differential equations. We demonstrate the method on experimental data from intravenous glucose tolerance tests performed on 19 normal glucose-tolerant subjects.

Adult↗

Molecular characterisation and evolution of the hemocyanin from the European spiny lobster, Palinurus elephas.

The hemocyanin of the European spiny lobster Palinurus elephas (synonym: Palinurus vulgaris) is a hexamer composed by four closely related but distinct subunits. We have obtained the full cDNA sequences of all four subunits, which cover 2275-2298 bp and encode for native polypeptides of 656 and 657 amino acids. The P. elephas hemocyanin subunits belong to the alpha-type of crustacean hemocyanins, whereas beta- and gamma-subunits are absent in this species. An unusual high ratio of non-synonymous versus synonymous nucleotide substitutions was observed, suggesting positive selection among subunits. Assuming a constant evolution rate, the P. elephas hemocyanin subunits emerged from a single hemocyanin gene around 25 million years ago. The alpha-type hemocyanins of P. elephas and the American spiny lobster Panulirus interruptus split around 100 million years ago. This is about five times older than the assumed divergence time of the species and suggests that the genera may have split with the formation of the Atlantic Ocean. The application of the Bayesian method for phylogenetic inference allows for the first time a solid reconstruction of the evolution of the decapod hemocyanins, showing that the beta-subunit types diverged first and that the crustacean pseudo-hemocyanins are associated with the gamma-type subunits.

Amino Acid Sequence↗

Uncertainty, neuromodulation, and attention.

Uncertainty in various forms plagues our interactions with the environment. In a Bayesian statistical framework, optimal inference and prediction, based on unreliable observations in changing contexts, require the representation and manipulation of different forms of uncertainty. We propose that the neuromodulators acetylcholine and norepinephrine play a major role in the brain's implementation of these uncertainty computations. Acetylcholine signals expected uncertainty, coming from known unreliability of predictive cues within a context. Norepinephrine signals unexpected uncertainty, as when unsignaled context switches produce strongly unexpected observations. These uncertainty signals interact to enable optimal inference and learning in noisy and changeable environments. This formulation is consistent with a wealth of physiological, pharmacological, and behavioral data implicating acetylcholine and norepinephrine in specific aspects of a range of cognitive processes. Moreover, the model suggests a class of attentional cueing tasks that involve both neuromodulators and shows how their interactions may be part-antagonistic, part-synergistic.

Acetylcholine↗