Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Simple and accurate estimation of ancestral protein sequences.

There are a variety of reasons to reconstruct the sequences of ancient proteins, but whatever the reason, the value of the reconstructed protein depends on the accuracy with which the ancient sequence is inferred. This study uses sequences simulated by a sequence-evolution simulation program that compares parsimony, maximum likelihood, and the Bayesian methods of inferring ancestral sequences and concludes that the Bayesian method, as implemented by MRBAYES 3.11, is preferred. Estimated ancestral sequences are of necessity the same length as the alignment on which the underlying phylogeny is based. A highly accurate method for correcting the estimated sequences is introduced, and it is shown that the correction permits inferring the sequences of ancient protein sequences with a very high degree of accuracy.

Base Sequence↗

An empirical Bayes approach to inferring large-scale gene association networks.

MOTIVATION: Genetic networks are often described statistically using graphical models (e.g. Bayesian networks). However, inferring the network structure offers a serious challenge in microarray analysis where the sample size is small compared to the number of considered genes. This renders many standard algorithms for graphical models inapplicable, and inferring genetic networks an 'ill-posed' inverse problem. METHODS: We introduce a novel framework for small-sample inference of graphical models from gene expression data. Specifically, we focus on the so-called graphical Gaussian models (GGMs) that are now frequently used to describe gene association networks and to detect conditionally dependent genes. Our new approach is based on (1) improved (regularized) small-sample point estimates of partial correlation, (2) an exact test of edge inclusion with adaptive estimation of the degree of freedom and (3) a heuristic network search based on false discovery rate multiple testing. Steps (2) and (3) correspond to an empirical Bayes estimate of the network topology. RESULTS: Using computer simulations, we investigate the sensitivity (power) and specificity (true negative rate) of the proposed framework to estimate GGMs from microarray data. This shows that it is possible to recover the true network topology with high accuracy even for small-sample datasets. Subsequently, we analyze gene expression data from a breast cancer tumor study and illustrate our approach by inferring a corresponding large-scale gene association network for 3883 genes.

Algorithms↗

Comparison of phylogenetic metrics of transmission in symptomatic and asymptomatic tuberculosis.

BACKGROUND: Understanding drivers of Mycobacterium tuberculosis (Mtb) transmission remains a critical challenge in high-burden settings. Tuberculosis control efforts traditionally target symptomatic individuals, yet the role of asymptomatic cases in sustaining transmission is increasing recognized. METHODS: We conducted a genomic and epidemiological analysis of Mtb isolates collected in Mato Grosso do Sul, Brazil, between 2008 and 2024. From 2017 to 2022, active case finding was performed in three of the state's largest prisons, whereby sputum was collected from individuals irrespective of symptoms and tested by GeneXpert and culture. We evaluated several metrics of recent transmission from symptomatic and asymptomatic individuals, including phylogenetic clustering, Time-scaled Haplotype Density (THD), Local Branching Index (LBI), and transmission probabilities inferred using the Bayesian Reconstruction and Evolutionary Analysis of Transmission Histories (BREATH). FINDINGS: We sequenced 2,362 Mtb strains, of which 3.5% (115/2,362) were resistant to at least one drug, and 0.6% (16/2,362) were multi-drug resistant. Most strains were lineage 4, and 78.2% of all isolates were part of a genomic cluster. Among 2,362 individuals with tuberculosis, 1,137 were incarcerated at the time of diagnosis. Among these, 505 were identified through active case finding: 277 had symptomatic disease and 228 had asymptomatic tuberculosis. There was no significant difference in phylogenetic clustering proportion (77% vs. 85%; p= 0.816), THD (median 0.50 vs. 0.39; p = 0.120), or LBI (median 0.00863 vs. 0.00871; p = 0.086) between symptomatic and asymptomatic individuals. Bayesian transmission trees revealed no significant difference in the number of secondary infections inferred from symptomatic compared with asymptomatic individuals (p = 0.56). These findings were consistent across genomic clusters and robust to model assumptions. INTERPRETATION: We identified no differences in transmission from symptomatic compared with asymptomatic individuals, using several genomic measures of transmission, underscoring the substantial contribution that asymptomatic tuberculosis makes to transmission at the population level.

Asymptomatic↗

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics↗

Inferring the root of a phylogenetic tree.

Phylogenetic trees can be rooted by a number of criteria. Here, we introduce a Bayesian method for inferring the root of a phylogenetic tree by using one of several criteria: the outgroup, molecular clock, and nonreversible model of DNA substitution. We perform simulation analyses to examine the relative ability of these three criteria to correctly identify the root of the tree. The outgroup and molecular clock criteria were best able to identify the root of the tree, whereas the nonreversible model was able to identify the root only when the substitution process was highly nonreversible. We also examined the performance of the criteria for a tree of four species for which the topology and root position are well supported. Results of the analyses of these data are consistent with the simulation results.

Bayes Theorem↗

The Bayesian controversy in animal breeding.

Frequentist and Bayesian approaches to scientific inference in animal breeding are discussed. Routine methods in animal breeding (selection index, BLUP, ML, REML) are presented under the hypotheses of both schools of inference, and their properties are examined in both cases. The Bayesian approach is discussed in cases in which prior information is available, prior information is available under certain hypotheses, prior information is vague, and there is no prior information. Bayesian prediction of genetic values and genetic parameters are presented. Finally, the frequentist and Bayesian approaches are compared from a theoretical and a practical point of view. Some problems for which Bayesian methods can be particularly useful are discussed. Both Bayesian and frequentist schools of inference are established, and now neither of them has operational difficulties, with the exception of some complex cases. There is software available to analyze a large variety of problems from either point of view. The choice of one school or the other should be related to whether there are solutions in one school that the other does not offer, to how easily the problems are solved, and to how comfortable scientists feel with the way they convey their results.

Animals↗

Neural coding: higher-order temporal patterns in the neurostatistics of cell assemblies.

Recent advances in the technology of multiunit recordings make it possible to test Hebb's hypothesis that neurons do not function in isolation but are organized in assemblies. This has created the need for statistical approaches to detecting the presence of spatiotemporal patterns of more than two neurons in neuron spike train data. We mention three possible measures for the presence of higher-order patterns of neural activation--coefficients of log-linear models, connected cumulants, and redundancies--and present arguments in favor of the coefficients of log-linear models. We present test statistics for detecting the presence of higher-order interactions in spike train data by parameterizing these interactions in terms of coefficients of log-linear models. We also present a Bayesian approach for inferring the existence or absence of interactions and estimating their strength. The two methods, the frequentist and the Bayesian one, are shown to be consistent in the sense that interactions that are detected by either method also tend to be detected by the other. A heuristic for the analysis of temporal patterns is also proposed. Finally, a Bayesian test is presented that establishes stochastic differences between recorded segments of data. The methods are applied to experimental data and synthetic data drawn from our statistical models. Our experimental data are drawn from multiunit recordings in the prefrontal cortex of behaving monkeys, the somatosensory cortex of anesthetized rats, and multiunit recordings in the visual cortex of behaving monkeys.

Action Potentials↗

Estimation of paratuberculosis prevalence in dairy cattle in a province of Korea using an enzyme-linked immunosorbent assay: application of Bayesian approach.

To draw inferences about the sensitivity and specificity of the newly developed ELISA test for bovine paratuberculosis (PTB) diagnosis and posterior distribution on the prevalence of PTB in a province of Korea, we applied Bayesian approach with Gibbs sampler to the data extracted from the prevalence study in 1999. The data were from a single test results without a designated gold test. The prevalence estimates for PTB in study population ranged 3.2-5.3% for conservative and 6.7-7.1% for liberal, depending on the priors used. The simulated specificities of the ELISA close to one another, ranging 84.7-90.6%, whereas the sensitivity was somewhat spread out depending largely on the priors with a range of 46.4-88.2%. Our findings indicate that the ELISA method appeared useful as a screening tool at a minimum level in comparison to other diagnostic tests available for this disease in terms of sensitivity. However, this advantage comes at a cost of having low specificity of the test.

Animals↗

A case for Bayesianism in clinical trials.

This paper describes a Bayesian approach to the design and analysis of clinical trials, and compares it with the frequentist approach. Both approaches address learning under uncertainty. But they are different in a variety of ways. The Bayesian approach is more flexible. For example, accumulating data from a clinical trial can be used to update Bayesian measures, independent of the design of the trial. Frequentist measures are tied to the design, and interim analyses must be planned for frequentist measures to have meaning. Its flexibility makes the Bayesian approach ideal for analysing data from clinical trials. In carrying out a Bayesian analysis for inferring treatment effect, information from the clinical trial and other sources can be combined and used explicitly in drawing conclusions. Bayesians and frequentists address making decisions very differently. For example, when choosing or modifying the design of a clinical trial, Bayesians use all available information, including that which comes from the trial itself. The ability to calculate predictive probabilities for future observations is a distinct advantage of the Bayesian approach to designing clinical trials and other decisions. An important difference between Bayesian and frequentist thinking is the role of randomization.

Bayes Theorem↗

Inferring gene networks from time series microarray data using dynamic Bayesian networks.

Dynamic Bayesian networks (DBNs) are considered as a promising model for inferring gene networks from time series microarray data. DBNs have overtaken Bayesian networks (BNs) as DBNs can construct cyclic regulations using time delay information. In this paper, a general framework for DBN modelling is outlined. Both discrete and continuous DBN models are constructed systematically and criteria for learning network structures are introduced from a Bayesian statistical viewpoint. This paper reviews the applications of DBNs over the past years. Real data applications for Saccharomyces cerevisiae time series gene expression data are also shown.

Algorithms↗

Bayesian mapping of QTL in outbred F2 families allowing inference about whether F0 grandparents are homozygous or heterozygous at QTL.

In this paper, we propose a new Bayesian method for QTL analysis in outbred F2 families based on Markov chain Monte Carlo (MCMC) estimation allowing inference about whether each of F0 founders (grandparents) is homozygous or heterozygous at QTL. This, in turn, allows us to select a model accurately explaining observations of phenotypes for F2 individuals. The proposed method performs the fitting a statistical model of the two possible QTL states in each F0 grandparent, that is, homozygous and heterozygous at QTL, and gives a posterior distribution for the QTL states in each F0 grandparent. We confine ourselves to the discrimination of two QTL states, homozygous or heterozygous, for each of the F0 grandparents without taking into consideration whether common alleles are shared by F0 grandparents. The statistical model includes allelic effects and dominance effects for each QTL. The number of parameters representing allelic effects and dominance effects is therefore changed depending on the QTL states. A Reversible Jump MCMC technique is used for transition between the models of different dimensions. The effectiveness of the proposed method was investigated using simulation experiments. It was practicable to estimate the QTL states of F0 grandparents as well as the number, the locations and the effects of QTL segregating in an outbred F2 family.

Alleles↗

New methods for detecting positive selection at single amino acid sites.

Inferring positive selection at single amino acid sites is of particular importance for studying evolutionary mechanisms of a protein. For this purpose, Suzuki and Gojobori (1999) developed a method (SG method) for comparing the rates of synonymous and nonsynonymous substitutions at each codon site in a protein-coding nucleotide sequence, using ancestral codons at interior nodes of the phylogenetic tree as inferred by the maximum parsimony method. In the SG method, however, selective neutrality of nucleotide substitutions cannot be tested at codon sites, where only termination codons are inferred at any interior node or the number of equally parsimonious inferences of ancestral codons at all interior nodes exceeds 10,000. Here I present a modified SG method which is free from these problems. Specifically, I use the distance-based Bayesian method for inferring the single most likely ancestral codon from 61 sense codons at each interior node. In the computer simulation and real data analysis, the modified SG method showed a higher overall efficiency of detecting positive selection than the original SG method, particularly at highly polymorphic codon sites. These results indicate that the modified SG method is useful for inferring positive selection at codon sites where neutrality cannot be tested by the original SG method. I also discuss that the p-distance is preferable to the number of synonymous substitutions for inferring the phylogenetic tree in the SG method, and present a maximum likelihood method for detecting positive selection at single amino acid sites, which produced reasonable results in the real data analysis.

Amino Acids↗

CYP11B2-CYP11B1 haplotypes associated with decreased 11 beta-hydroxylase activity.

Reduced adrenal 11 beta-hydroxylation has been associated with an aldosterone synthase (CYP11B2) polymorphism. The 11 beta-hydroxylase gene (CYP11B1) lies close to CYP11B2. We hypothesize that a molecular variant in CYP11B2 is in linkage disequilibrium (LD) with a key quantitative trait in CYP11B1 determining this phenotype. Polymorphisms and inferred haplotypes at CYP11B loci were studied in two independent populations from Europe (n = 100) and South America (n = 99). The latter underwent detailed hormonal studies. LD was estimated by alternative Bayesian methods for inferring the extent of LD when haplotypes at different loci are inferred. Population differences in single nucleotide polymorphisms were modest, indicating the stability of both genes across populations. Using five of nine potentially informative loci at CYP11B sites with allele frequency greater than 0.1, two major contrasting haplotypes, CwtCG and TconvGTA, were found. In both populations the CwtCG haplotype accounted for 44% and the TconvGTA for 32% of subjects. Haplotype distribution did not differ between Europeans and South Americans (chi(2) = 2.81; P = 0.09). In vivo 11 beta-hydroxylase activity, estimated from urinary steroid profiling, was lower in subjects with an increased aldosterone to renin ratio or with the TconvGTA haplotype. These findings indicate that genotypes at the CYP11B locus are in strong LD and that identified haplotypes predict 11 beta-hydroxylase activity.

Base Sequence↗

Computational inference of neural information flow networks.

Determining how information flows along anatomical brain pathways is a fundamental requirement for understanding how animals perceive their environments, learn, and behave. Attempts to reveal such neural information flow have been made using linear computational methods, but neural interactions are known to be nonlinear. Here, we demonstrate that a dynamic Bayesian network (DBN) inference algorithm we originally developed to infer nonlinear transcriptional regulatory networks from gene expression data collected with microarrays is also successful at inferring nonlinear neural information flow networks from electrophysiology data collected with microelectrode arrays. The inferred networks we recover from the songbird auditory pathway are correctly restricted to a subset of known anatomical paths, are consistent with timing of the system, and reveal both the importance of reciprocal feedback in auditory processing and greater information flow to higher-order auditory areas when birds hear natural as opposed to synthetic sounds. A linear method applied to the same data incorrectly produces networks with information flow to non-neural tissue and over paths known not to exist. To our knowledge, this study represents the first biologically validated demonstration of an algorithm to successfully infer neural information flow networks.

Action Potentials↗

Molecular systematics of Salmonidae: combined nuclear data yields a robust phylogeny.

The phylogeny of salmonid fishes has been the focus of intensive study for many years, but some of the most important relationships within this group remain unclear. We used 269 Genbank sequences of mitochondrial DNA (from 16 genes) and nuclear DNA (from nine genes) to infer phylogenies for 30 species of salmonids. We used maximum parsimony and maximum likelihood to analyze each gene separately, the mtDNA data combined, the nuclear data combined, and all of the data together. The phylogeny with the best overall resolution and support from bootstrapping and Bayesian analyses was inferred from the combined nuclear DNA data set, for which the different genes reinforced and complemented one another to a considerable degree. Addition of the mitochondrial DNA degraded the phylogenetic signal, apparently as a result of saturation, hybridization, selection, or some combination of these processes. By the nuclear-DNA phylogeny: (1) (Hucho hucho, Brachymystax lenok) form the sister group to (Salmo, Salvelinus, Oncorhynchus, H. perryi); (2) Salmo is the sister-group to (Oncorhynchus, Salvelinus); (3) Salvelinus is the sister-group to Oncorhynchus; and (4) Oncorhynchus masou forms a monophyletic group with O. mykiss and O. clarki, with these three taxa constituting the sister-group to the five other Oncorhynchus species. Species-level relationships within Oncorhynchus and Salvelinus were well supported by bootstrap levels and Bayesian analyses. These findings have important implications for understanding the evolution of behavior, ecology and life-history in Salmonidae.

Animals↗

A bayesian statistical algorithm for RNA secondary structure prediction.

A Bayesian approach for predicting RNA secondary structure that addresses the following three open issues is described: (1) the need for a representation of the full ensemble of probable structures; (2) the need to specify a fixed set of energy parameters; (3) the desire to make statistical inferences on all variables in the problem. It has recently been shown that Bayesian inference can be employed to relax or eliminate the need to specify the parameters of bioinformatics recursive algorithms and to give a statistical representation of the full ensemble of probable solutions with the incorporation of uncertainty in parameter values. In this paper, we make an initial exploration of these potential advantages of the Bayesian approach. We present a Bayesian algorithm that is based on stacking energy rules but relaxes the need to specify the parameters. The algorithm returns the exact posterior distribution of the number of destabilizing loops, stacking energy matrices, and secondary structures. The algorithm generates statistically representative structures from the full ensemble of probable secondary structures in exact proportion to the posterior probabilities. Once the forward recursions for the algorithm are completed, the backward recursive sampling executes in O(n) time, providing a very efficient approach for generating representative structures. We demonstrate the utility of the Bayesian approach with several tRNA sequences. The potential of the approach for predicting RNA secondary structures and presenting alternative structures is illustrated with applications to the Escherichia coli tRNA(Ala) sequence and the Xenopus laevis oocyte 5S rRNA sequence.

Algorithms↗

Selective integration of multiple biological data for supervised network inference.

MOTIVATION: Inferring networks of proteins from biological data is a central issue of computational biology. Most network inference methods, including Bayesian networks, take unsupervised approaches in which the network is totally unknown in the beginning, and all the edges have to be predicted. A more realistic supervised framework, proposed recently, assumes that a substantial part of the network is known. We propose a new kernel-based method for supervised graph inference based on multiple types of biological datasets such as gene expression, phylogenetic profiles and amino acid sequences. Notably, our method assigns a weight to each type of dataset and thereby selects informative ones. Data selection is useful for reducing data collection costs. For example, when a similar network inference problem must be solved for other organisms, the dataset excluded by our algorithm need not be collected. RESULTS: First, we formulate supervised network inference as a kernel matrix completion problem, where the inference of edges boils down to estimation of missing entries of a kernel matrix. Then, an expectation-maximization algorithm is proposed to simultaneously infer the missing entries of the kernel matrix and the weights of multiple datasets. By introducing the weights, we can integrate multiple datasets selectively and thereby exclude irrelevant and noisy datasets. Our approach is favorably tested in two biological networks: a metabolic network and a protein interaction network. AVAILABILITY: Software is available on request.

Algorithms↗

A model to estimate the optimal sample size for microbiological surveys.

Estimating optimal sample size for microbiological surveys is a challenge for laboratory managers. When insufficient sampling is conducted, biased inferences are likely; however, when excessive sampling is conducted valuable laboratory resources are wasted. This report presents a statistical model for the estimation of the sample size appropriate for the accurate identification of the bacterial subtypes of interest in a specimen. This applied model for microbiology laboratory use is based on a Bayesian mode of inference, which combines two inputs: (ii) a prespecified estimate, or prior distribution statement, based on available scientific knowledge and (ii) observed data. The specific inputs for the model are a prior distribution statement of the number of strains per specimen provided by an informed microbiologist and data from a microbiological survey indicating the number of strains per specimen. The model output is an updated probability distribution of strains per specimen, which can be used to estimate the probability of observing all strains present according to the number of colonies that are sampled. In this report two scenarios that illustrate the use of the model to estimate bacterial colony sample size requirements are presented. In the first scenario, bacterial colony sample size is estimated to correctly identify Campylobacter amplified restriction fragment length polymorphism types on broiler carcasses. The second scenario estimates bacterial colony sample size to correctly identify Salmonella enterica serotype Enteritidis phage types in fecal drag swabs from egg-laying poultry flocks. An advantage of the model is that as updated inputs from ongoing surveys are incorporated into the model, increasingly precise sample size estimates are likely to be made.

Animals↗