Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Bayesian estimation of relaxation times T(1) in MR images of irradiated Fricke-agarose gels.

The authors present a novel method for processing T(1)-weighted images acquired with Inversion-Recovery (IR) sequence. The method, developed within the Bayesian framework, takes into account a priori knowledge about the spatial regularity of the parameters to be estimated. Inference is drawn by means of Markov Chains Monte Carlo algorithms. The method has been applied to the processing of IR images from irradiated Fricke-agarose gels, proposed in the past as relative dosimeter to verify radiotherapeutic treatment planning systems. Comparison with results obtained from a standard approach shows that signal-to noise ratio (SNR) is strongly enhanced when the estimation of the longitudinal relaxation rate (R1) is performed with the newly proposed statistical approach. Furthermore, the method allows the use of more complex models of the signal. Finally, an appreciable reduction of total acquisition time can be obtained due to the possibility of using a reduced number of images. The method can also be applied to T(1) mapping of other systems.

Bayes Theorem↗

Towards improved fine-mapping of candidate causal variants.

Fine-mapping in genome-wide association studies aims to identify potentially causal genetic variants among a set of candidate variants that are often highly correlated with each other owing to linkage disequilibrium. A variety of statistical approaches are used in fine-mapping, almost all of which are based on a multiple regression framework to model the relationship between genotype and phenotype, while accommodating specific assumptions about the distribution of variant effect sizes and using different inference algorithms. Owing to their modelling flexibility and the ease of making inferential statements, these approaches are predominantly Bayesian in nature. Recently, these approaches have been improved by refining modelling assumptions, integrating additional information, accommodating summary statistics, and developing scalable computational algorithms that improve computation efficiency and fine-mapping resolution.

Humans↗

Small-sample inference for incomplete longitudinal data with truncation and censoring in tumor xenograft models.

In cancer drug development, demonstrating activity in xenograft models, where mice are grafted with human cancer cells, is an important step in bringing a promising compound to humans. A key outcome variable is the tumor volume measured in a given period of time for groups of mice given different doses of a single or combination anticancer regimen. However, a mouse may die before the end of a study or may be sacrificed when its tumor volume quadruples, and its tumor may be suppressed for some time and then grow back. Thus, incomplete repeated measurements arise. The incompleteness or missingness is also caused by drastic tumor shrinkage (<0.01 cm3) or random truncation. Because of the small sample sizes in these models, asymptotic inferences are usually not appropriate. We propose two parametric test procedures based on the EM algorithm and the Bayesian method to compare treatment effects among different groups while accounting for informative censoring. A real xenograft study on a new antitumor agent, temozolomide, combined with irinotecan is analyzed using the proposed methods.

Algorithms↗

Genetic analysis of growth curves using the SAEM algorithm.

The analysis of nonlinear function-valued characters is very important in genetic studies, especially for growth traits of agricultural and laboratory species. Inference in nonlinear mixed effects models is, however, quite complex and is usually based on likelihood approximations or Bayesian methods. The aim of this paper was to present an efficient stochastic EM procedure, namely the SAEM algorithm, which is much faster to converge than the classical Monte Carlo EM algorithm and Bayesian estimation procedures, does not require specification of prior distributions and is quite robust to the choice of starting values. The key idea is to recycle the simulated values from one iteration to the next in the EM algorithm, which considerably accelerates the convergence. A simulation study is presented which confirms the advantages of this estimation procedure in the case of a genetic analysis. The SAEM algorithm was applied to real data sets on growth measurements in beef cattle and in chickens. The proposed estimation procedure, as the classical Monte Carlo EM algorithm, provides significance tests on the parameters and likelihood based model comparison criteria to compare the nonlinear models with other longitudinal methods.

Algorithms↗

Identification and genetic validation of potential therapeutic targets for pulmonary hypertension through multi-omics causal inference.

Pulmonary hypertension (PH) underscores the urgent need for novel therapeutic targets. This study aimed to employ a proteome-wide Mendelian randomization (MR) approach to systematically identify circulating proteins causally associated with PH, thereby providing genetically validated candidate targets for drug development. We adopted a 2-sample MR design, integrating large-scale plasma proteomic quantitative trait loci (pQTL) data (encompassing 4148 proteins) and summary statistics from a large-scale PH genome-wide association study (2047 cases, 8301 controls). Candidate targets were screened through a multilayered analytical pipeline comprising proteomic MR, transcriptomic MR, and summary-data-based Mendelian randomization. The ultimately identified MR-Identified Causal Candidate Targets (MR-ICTs) underwent rigorous Bayesian colocalization analysis, followed by biological characterization through functional enrichment analysis, single-cell transcriptomics, and phenome-wide association studies. Through robust genetic causal inference, this study provides that circulating proteins such as LYZ, GREM2, NID1, and PF4V1 play causal roles in PH pathogenesis. These findings offer a set of rigorously genetically validated, high-priority therapeutic targets for developing novel PH treatments, specifically addressing key pathological mechanisms such as innate immunity, BMP signaling pathway dysregulation, and platelet activation. Our multi-dimensional analysis ultimately identified 6 MR-ICTs causally associated with PH. Notably, the causal associations for lysozyme C (LYZ), gremlin-2 (GREM2), nidogen-1 (NID1), and platelet factor 4 variant 1 (PF4V1) were stringently validated by Bayesian colocalization analysis (posterior probability for hypothesis 4 [PPH4], indicating a shared causal variant, > 0.99). Functional enrichment analysis revealed significant involvement of these targets in immune response and TGF-&#x3b2; signaling pathways. Single-cell analysis further elucidated their cell-type-specific expression, with LYZ predominantly expressed in monocytes and PF4V1 almost exclusively in platelets.

Hypertension, Pulmonary↗

Molecular phylogenetics and biogeography of Lepus in Eastern Asia based on mitochondrial DNA sequences.

In spite of several classification attempts among taxa of the genus Lepus, phylogenetic relationships still remain poorly understood. Here, we present molecular genetic evidence that may resolve some of the current incongruities in the phylogeny of the leporids. The complete mitochondrial cytb, 12S genes, and parts of ND4 and control region fragments were sequenced to examine phylogenetic relationships among Chinese hare taxa and other leporids throughout the World using maximum parsimony, maximum likelihood, and Bayesian phylogenetic reconstruction approaches. Using reconstructed phylogenies, we observed that the Chinese hare is not a single monophyletic group as originally thought. Instead, the data infers that the genus Lepus is monophyletic with three unique species groups: North American, Eurasian, and African. Ancestral area analysis indicated that ancestral Lepus arose in North America and then dispersed into Eurasia via the Bering Land Bridge eventually extending to Africa. Brooks Parsimony analysis showed that dispersal events followed by subsequent speciation have occurred in other geographic areas as well and resulted in the rapid radiation and speciation of Lepus. A Bayesian relaxed molecular clock approach based on the continuous autocorrelation of evolutionary rates along branches estimated the divergence time between the three major groups within Lepus. The genus appears to have arisen approximately 10.76 MYA (+/-0.86 MYA), with most speciation events occurring during the Pliocene epoch (5.65+/-1.15 MYA approximately 1.12 +/- 0.47 MYA).

Animals↗

Prediction of splice sites with dependency graphs and their expanded bayesian networks.

MOTIVATION: Owing to the complete sequencing of human and many other genomes, huge amounts of DNA sequence data have been accumulated. In bioinformatics, an important issue is how to predict the complete structure of genes from the genomic DNA sequence, especially the human genome. A crucial part in the gene structure prediction is to determine the precise exon-intron boundaries, i.e. the splice sites, in the coding region. RESULTS: We have developed a dependency graph model to fully capture the intrinsic interdependency between base positions in a splice site. The establishment of dependency between two position is based on a chi2-test from known sample data. To facilitate statistical inference, we have expanded the dependency graph (which is usually a graph with cycles that make probabilistic reasoning very difficult, if not impossible) into a Bayesian network (which is a directed acyclic graph that facilitates statistical reasoning). When compared with the existing models such as weight matrix model, weight array model, maximal dependence decomposition, Cai et al.'s tree model as well as the less-studied second-order and third-order Markov chain models, the expanded Bayesian networks from our dependency graph models perform the best in nearly all the cases studied. AVAILABILITY: Software (a program called DGSplicer) and datasets used are available at http://csrl.ee.nthu.edu.tw/bioinf/ CONTACT: cclu@ee.nthu.edu.tw.

Bayes Theorem↗

Bayesian synthesis of a pathogen growth model: Listeria monocytogenes under competition.

The Bayesian synthesis method is applied to data from two studies of Listeria monocytogenes grown in broth monocultures to draw inferences about the joint distribution of two Baranyi growth model parameters-lag time and maximum specific growth rate. The resultant joint distribution is then combined with prior distributions for the initial and maximum pathogen density parameters under competitive growth conditions. Finally, the pathogen growth model is updated using the Sampling/Importance Resampling (SIR) algorithm with data on L. monocytogenes growth in competition with natural microflora in fish. Although the latter data provide no information on the stationary phase to directly estimate the maximum pathogen density parameter, combining them with relevant prior information provides a means to characterize L. monocytogenes growth in a food with mixed microbial populations. Based on a specified tolerance for L. monocytogenes growth, the updated model provides a storage time limit for fish held at 5 degrees C, pH 6.8, 43% CO(2), 57% N(2).

Animals↗

Implementation of automated signal generation in pharmacovigilance using a knowledge-based approach.

Automated signal generation is a growing field in pharmacovigilance that relies on data mining of huge spontaneous reporting systems for detecting unknown adverse drug reactions (ADR). Previous implementations of quantitative techniques did not take into account issues related to the medical dictionary for regulatory activities (MedDRA) terminology used for coding ADRs. MedDRA is a first generation terminology lacking formal definitions; grouping of similar medical conditions is not accurate due to taxonomic limitations. Our objective was to build a data-mining tool that improves signal detection algorithms by performing terminological reasoning on MedDRA codes described with the DAML+OIL description logic. We propose the PharmaMiner tool that implements quantitative techniques based on underlying statistical and bayesian models. It is a JAVA application displaying results in tabular format and performing terminological reasoning with the Racer inference engine. The mean frequency of drug-adverse effect associations in the French database was 2.66. Subsumption reasoning based on MedDRA taxonomical hierarchy produced a mean number of occurrence of 2.92 versus 3.63 (p < 0.001) obtained with a combined technique using subsumption and approximate matching reasoning based on the ontological structure. Semantic integration of terminological systems with data mining methods is a promising technique for improving machine learning in medical databases.

Adverse Drug Reaction Reporting Systems↗

Molecular evidence for the monophyly of East Asian groups of Cyprinidae (Teleostei: Cypriniformes) derived from the nuclear recombination activating gene 2 sequences.

The family Cyprinidae is one of the largest families of fishes in the world and a well-known component of the East Asian freshwater fish fauna. However, the phylogenetic relationships among cyprinids are still poorly understood despite much effort paid on the cyprinid molecular phylogenetics. Original nucleotide sequence data of the nuclear recombination activating gene 2 were collected from 109 cyprinid species and four non-cyprinid cypriniform outgroup taxa and used to infer the cyprinid phylogenetic relationships and to estimate node divergence times. Phylogenetic reconstructions using maximum parsimony, maximum likelihood, and Bayesian analysis retrieved the same clades, only branching order within these clades varied slightly between trees. Although the morphological diversity is remarkable, the endemic cyprinid taxa in East Asia emerged as a monophyletic clade referred to as Xenocypridini. The monophyly for the subfamilies including Cyprininae and Leuciscinae, as well as the tribes including Labeonini, Gobionini, Acheilognathini, and Leuciscini, was also well resolved with high nodal support. Analysis of the RAG2 gene supported the following cyprinid molecular phylogeny: the Danioninae is the most basal subfamily within the family Cyprinidae and the Cyprininae is the sister group of the Leuciscinae. The divergence times were estimated for the nodes corresponding to the principal clades within the Cyprinidae. The family Cyprinidae appears to have originated in the mid-Eocene in Asia, with the cladogenic event of the key basal group Danioninae occurring in the early Oligocene (about 31-30 MYA), and the origins of the two subfamilies, Cyprininae and Leuciscinae, occurring in the mid-Oligocene (around 26 MYA).

Animals↗

Molecular analysis reveals tighter social regulation of immigration in patrilocal populations than in matrilocal populations.

Human social organization can deeply affect levels of genetic diversity. This fact implies that genetic information can be used to study social structures, which is the basis of ethnogenetics. Recently, methods have been developed to extract this information from genetic data gathered from subdivided populations that have gone through recent spatial expansions, which is typical of most human populations. Here, we perform a Bayesian analysis of mitochondrial and Y chromosome diversity in three matrilocal and three patrilocal groups from northern Thailand to infer the number of males and females arriving in these populations each generation and to estimate the age of their range expansion. We find that the number of male immigrants is 8 times smaller in patrilocal populations than in matrilocal populations, whereas women move 2.5 times more in patrilocal populations than in matrilocal populations. In addition to providing genetic quantification of sex-specific dispersal rates in human populations, we show that although men and women are exchanged at a similar rate between matrilocal populations, there are far fewer men than women moving into patrilocal populations. This finding is compatible with the hypothesis that men are strictly controlling male immigration and promoting female immigration in patrilocal populations and that immigration is much less regulated in matrilocal populations.

Bayes Theorem↗

Phylogeny of eusocial Lasioglossum reveals multiple losses of eusociality within a primitively eusocial clade of bees (Hymenoptera: Halictidae).

We performed a phylogenetic analysis of the species, species groups, and subgenera within the predominantly eusocial lineage of Lasioglossum (the Hemihalictus series) based on three protein coding genes: mitochondrial cytochrome oxidase I, nuclear elongation factor 1alpha and long-wavelength rhodopsin. The entire data set consisted of 3421 aligned nucleotide sites, 854 of which were parsimony informative. Analyses by equal weights parsimony, maximum likelihood, and Bayesian methods yielded good resolution among the 53 taxa/populations, with strong bootstrap support and high posterior probabilities for most nodes. There was no significant incongruence among genes, and parsimony, maximum likelihood, and Bayesian methods yielded congruent results. We mapped social behavior onto the resulting tree for 42 of the taxa/populations to infer the likely history of social evolution within Lasioglossum. Our results indicate that eusociality had a single origin within Lasioglossum. Within the predominantly eusocial clade, however, there have been multiple (six) reversals from eusociality to solitary nesting, social polymorphism, or social parasitism, suggesting that these reversals may be more common in primitively eusocial Hymenoptera than previously anticipated. Our results support the view that eusociality is hard to evolve but easily lost. This conclusion is potentially important for understanding the early evolution of the advanced eusocial insects, such as ants, termites, and corbiculate bees.

Animals↗

Reconstruction of gene networks using Bayesian learning and manipulation experiments.

MOTIVATION: The analysis of high-throughput experimental data, for example from microarray experiments, is currently seen as a promising way of finding regulatory relationships between genes. Bayesian networks have been suggested for learning gene regulatory networks from observational data. Not all causal relationships can be inferred from correlation data alone. Often several equivalent but different directed graphs explain the data equally well. Intervention experiments where genes are manipulated can help to narrow down the range of possible networks. RESULTS: We describe an active learning algorithm that suggests an optimized sequence of intervention experiments. Simulation experiments show that our selection scheme is better than an unguided choice of interventions in learning the correct network and compares favorably in running time and results with methods based on value of information calculations.

Algorithms↗

Binomial regression with misclassification.

Motivated by a study of human papillomavirus infection in women, we present a Bayesian binomial regression analysis in which the response is subject to an unconstrained misclassification process. Our iterative approach provides inferences for the parameters that describe the relationships of the covariates with the response and for the misclassification probabilities. Furthermore, our approach applies to any meaningful generalized linear model, making model selection possible. Finally, it is straightforward to extend it to multinomial settings.

Bayes Theorem↗

A Bayesian hierarchical approach for combining case-control and prospective studies.

Motivated by the absolute risk predictions required in medical decision making and patient counseling, we propose an approach for the combined analysis of case-control and prospective studies of disease risk factors. The approach is hierarchical to account for parameter heterogeneity among studies and among sampling units of the same study. It is based on modeling the retrospective distribution of the covariates given the disease outcome, a strategy that greatly simplifies both the combination of prospective and retrospective studies and the computation of Bayesian predictions in the hierarchical case-control context. Retrospective modeling differentiates our approach from most current strategies for inference on risk factors, which are based on the assumption of a specific prospective model. To ensure modeling flexibility, we propose using a mixture model for the retrospective distributions of the covariates. This leads to a general nonlinear regression family for the implied prospective likelihood. After introducing and motivating our proposal, we present simple results that highlight its relationship with existing approaches, develop Markov chain Monte Carlo methods for inference and prediction, and present an illustration using ovarian cancer data.

Bayes Theorem↗

A hierarchical aggregate data model with spatially correlated disease rates.

The aggregate data study design (Prentice and Sheppard, 1995, Biometrika 82, 113-125) estimates individual-level exposure effects by regressing population-based disease rates on covariate data from survey samples in each population group. In this work, we further develop the aggregate data model to allow for residual spatial correlation among disease rates across populations. Geographical variation that is not explained by model predictors and has a spatial component often arises in studies of rare chronic diseases, such as breast cancer. We combine the aggregate and Bayesian disease-mapping models to provide an intuitive approach to the modeling of spatial effects while drawing correct inference regarding the exposure effect. Based on the results of simulation studies, we suggest guidelines for use of the proposed model.

Bayes Theorem↗

On the origins of medfly invasion and expansion in Australia.

As a result of their rapid expansion and large larval host range, true fruit flies are among the world's most important agricultural pest species. Among them, Ceratitis capitata has become a model organism for studies on colonization and invasion processes. The genetic aspects of the medfly invasion process have already been analysed throughout its range, with the exception of Australia. Bioinvasion into Australia is an old event: medfly were first captured in Australia in 1895, near Perth. After briefly appearing in Tasmania and the eastern states of mainland Australia, medfly had disappeared from these areas by the 1940s. Currently, they are confined to the western coastal region. South Australia seems to be protected from medfly infestations both by the presence of an inhospitable barrier separating it from the west and by the limited number of transport routes. However, numerous medfly outbreaks have occurred since 1946, mainly near Adelaide. Allele frequency data at 10 simple sequence repeat loci were used to study the genetic structure of Australian medflies, to infer the historical pattern of invasion and the origin of the recent outbreaks. The combination of phylogeographical analysis and Bayesian tests showed that colonization of Australia was a secondary colonization event from the Mediterranean basin and that Australian medflies were unlikely to be the source for the initial Hawaiian invasion. Within Australia, the Perth area acted as the core range and was the source for medfly bioinvasion in both Western and South Australia. Incipient differentiation, as a result of habitat fragmentation, was detected in some localized areas at the periphery of the core range.

Animals↗

Who uses base rates and P(D/approximately H)? An analysis of individual differences.

In two experiments, involving over 900 subjects, we examined the cognitive correlates of the tendency to view P(D/approximately H) and base rate information as relevant to probability assessment. We found that individuals who viewed P(D/approximately H) as relevant in a selection task and who used it to make the proper Bayesian adjustment in a probability assessment task scored higher on tests of cognitive ability and were better deductive and inductive reasoners. They were less biased by prior beliefs and more data-driven on a covariation assessment task. In contrast, individuals who thought that base rates were relevant did not display better reasoning skill or higher cognitive ability. Our results parallel disputes about the normative status of various components of the Bayesian formula in interesting ways. It is argued that patterns of covariance among reasoning tasks may have implications for inferences about what individuals are trying to optimize in a rational analysis (J. R. Anderson, 1990, 1991).

Adolescent↗