Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Trying to be precise about vagueness.

A previous investigation by Lambert et al., which used computer simulation to examine the influence of choice of prior distribution on inferences from Bayesian random effects meta-analysis, is critically examined from a number of viewpoints. The practical example used is shown to be problematic. The various prior distributions are shown to be unreasonable in terms of what they imply about the joint distribution of the overall treatment effect and the random effects variance. An alternative form of prior distribution is tentatively proposed. Finally, some practical recommendations are made that stress the value both of fixed effect analyses and of frequentist approaches as well as various diagnostic investigations.

Bayes Theorem↗

Population toxicokinetics of tetrachloroethylene.

In assessing the distribution and metabolism of toxic compounds in the body, measurements are not always feasible for ethical or technical reasons. Computer modeling offers a reasonable alternative, but the variability and complexity of biological systems pose unique challenges in model building and adjustment. Recent tools from population pharmacokinetics, Bayesian statistical inference, and physiological modeling can be brought together to solve these problems. As an example, we modeled the distribution and metabolism of tetrachloroethylene (PERC) in humans. We derive statistical distributions for the parameters of a physiological model of PERC, on the basis of data from Monster et al. (1979). The model adequately fits both prior physiological information and experimental data. An estimate of the relationship between PERC exposure and fraction metabolized is obtained. Our median population estimate for the fraction of inhaled tetrachloroethylene that is metabolized, at exposure levels exceeding current occupational standards, is 1.5% [95% confidence interval (0.52%, 4.1%)]. At levels approaching ambient inhalation exposure (0.001 ppm), the median estimate of the fraction metabolized is much higher, at 36% [95% confidence interval (15%, 58%)]. This disproportionality should be taken into account when deriving safe exposure limits for tetrachloroethylene and deserves to be verified by further experiments.

Administration, Inhalation↗

Core passive and facultative mTOR-mediated mechanisms coordinate mammalian protein synthesis and decay.

The maintenance of cellular homeostasis requires tight regulation of proteome concentration and composition. To achieve this, protein production and elimination must be robustly coordinated. However, the mechanistic basis of this coordination remains unclear. Here, we address this question using quantitative live-cell imaging, computational modeling, transcriptomics, and proteomics approaches. We found that protein decay rates systematically adapt to global alterations of protein synthesis rates. This adaptation is driven by a core passive mechanism supplemented by facultative changes in mechanistic/mammalian target of rapamycin (mTOR) signaling. Passive adaptation hinges on changes in the production rate of the machinery governing protein decay and allows for partial maintenance of the cellular proteome. Sustained changes in mTOR signaling provide an additional layer of adaptation unique to naive pluripotent stem cells, allowing for near-perfect maintenance of proteome composition. Our work unravels the mechanisms protecting the integrity of mammalian proteomes upon variations in protein synthesis rates. A record of this paper's transparent peer review process is included in the supplemental information.

TOR Serine-Threonine Kinases↗

Analyzing infant mortality with geoadditive categorical regression models: a case study for Nigeria.

In this paper, we analyze infant mortality in Nigeria based on the data set from the 1999 Nigeria Demographic and Health Survey (NDHS). We investigate spatial patterns at a highly disaggregated level of Nigerian states and consider non-linear effects of mother's age at birth. Time to the occurrence of a child's death can intuitively be considered to be categorical in nature and the determinants of a child's death may differ in different age groups. Thus, it may be desirable to investigate separately the death of a child in the first month and in the remaining 11 months of the first year of life. To avoid selection bias, the data set used for this case study is based on information on children who were born 12 months preceding the survey. Inference is Bayesian and is based on Markov chain Monte Carlo (MCMC) techniques. We find that spatial variation and the determinants of death indeed differ considerably for the two age groups considered.

Adolescent↗

SARS associated coronavirus has a recombinant polymerase and coronaviruses have a history of host-shifting.

The sudden appearance and potential lethality of severe acute respiratory syndrome associated coronavirus (SARS-CoV) in humans has focused attention on understanding its origins. Here, we assess phylogenetic relationships for the SARS-CoV lineage as well as the history of host-species shifts for SARS-CoV and other coronaviruses. We used a Bayesian phylogenetic inference approach with sliding window analyses of three SARS-CoV proteins: RNA dependent RNA polymerase (RDRP), nucleocapsid (N) and spike (S). Conservation of RDRP allowed us to use a set of Arteriviridae taxa to root the Coronaviridae phylogeny. We found strong evidence for a recombination breakpoint within SARS-CoV RDRP, based on different, well supported trees for a 5' fragment (supporting SARS-CoV as sister to a clade including all other coronaviruses) and a 3' fragment (supporting SARS-CoV as sister to group three avian coronaviruses). These different topologies are statistically significant: the optimal 5' tree could be rejected for the 3' region, and the optimal 3' tree could be rejected for the 5' region. We did not find statistical evidence for recombination in analyses of N and S, as there is little signal to differentiate among alternative trees. Comparison of phylogenetic trees for 11 known host-species and 36 coronaviruses, representing coronavirus groups 1-3 and SARS-CoV, based on N showed statistical incongruence indicating multiple host-species shifts for coronaviruses. Inference of host-species associations is highly sensitive to sampling and must be considered cautiously. However, current sampling suggests host-species shifts between mouse and rat, chicken and turkey, mammals and manx shearwater, and humans and other mammals. The sister relationship between avian coronaviruses and the 3' RDRP fragment of SARS-CoV suggests an additional host-species shift. Demonstration of recombination in the SARS-CoV lineage indicates its potential for rapid unpredictable change, a potentially important challenge for public health management and for drug and vaccine development.

Animals↗

Molecular epidemiology and phylogeographic architecture of oncogenic intracellular bacteria in cervical cancer patients across Northern China.

BACKGROUND: Oncogenic intracellular bacteria, including Chlamydia trachomatis, Mycoplasma genitalium, and Fusobacterium nucleatum, have emerged as significant contributors to cervical carcinogenesis. Despite growing interest in microbial oncology, the molecular epidemiological landscape and phylogeographic distribution of these pathogens in Northern China remain poorly characterized. This study aimed to determine the prevalence, co-infection patterns, genotypic diversity, and spatial phylogeographic clustering of oncogenic intracellular bacteria among cervical cancer patients across five provinces of Northern China. METHODS: A cross-sectional, multi-center study was conducted between March 2022 and November 2024 across Shaanxi, Heilongjiang, Beijing, Shandong, and Inner Mongolia. Cervical swab specimens were collected from 1247 confirmed cervical cancer patients. Pathogen detection was performed using multiplex real-time polymerase chain reaction, 16S rRNA gene amplicon sequencing, and whole-genome sequencing. Phylogeographic analyses employed maximum likelihood and Bayesian evolutionary inference frameworks. Statistical analyses included multivariate logistic regression and geographic information system-based spatial clustering. RESULTS: The overall prevalence of at least one oncogenic intracellular bacterium was 68.3% (n&#xa0;=&#xa0;852). Chlamydia trachomatis was the most prevalent pathogen detected in 41.2% of participants. Co-infection with two or more bacteria was identified in 29.7% of cases and was independently associated with advanced-stage cervical cancer (adjusted odds ratio&#xa0;=&#xa0;2.87; 95% confidence interval: 1.94 to 4.23; p&#xa0;<&#xa0;0.001). Phylogeographic analysis revealed three distinct molecular clades with evidence of bidirectional gene flow between Shaanxi and Heilongjiang. Whole-genome sequencing identified 14 novel virulence gene variants not previously characterized in Chinese clinical isolates. CONCLUSIONS: Oncogenic intracellular bacteria are highly prevalent and genotypically diverse among cervical cancer patients in Northern China. The identified phylogeographic clustering and novel virulence variants have direct implications for regional screening programs, targeted antimicrobial strategies, and the development of region-specific molecular diagnostic panels.

Cervical cancer↗

Immune cell-specific genetic architecture of Alzheimer's disease revealed by multi-omics analysis for therapeutic target discovery and prioritization.

Alzheimer's disease (AD) is a multifactorial neurodegenerative condition in which accumulating genetic and molecular evidence implicates dysregulation of peripheral immune processes in disease pathogenesis. Nevertheless, the contribution of distinct peripheral immune cell subsets and associated gene regulatory landscapes to AD risk remains incompletely defined. To address this gap, we integrated single-cell expression quantitative trait loci (sc&#x2011;eQTL) data from the OneK1K cohort with AD GWAS summary statistics. We systematically interrogated immune cell-specific genes for their contributions to AD risk by integrating genetic causal inference with Bayesian colocalization analyses, and identified 24 eGenes that passed both the MR significance threshold (P&#x2009;<&#x2009;0.05) and the criterion for strong shared genetic signals (PP.H4&#x2009;>&#x2009;0.8). Notable candidates included GATS, HLA-DOB, HLA-DQA1, PM20D1, and others, with each gene demonstrating a cell-type-specific association restricted to its corresponding immune cell type, such as monocytes, CD8&#x2009;+&#x2009;T cells, or B cells. Independent peripheral blood single-cell transcriptomic data further supported disease-associated shifts in cell-type-specific expression patterns in AD. Phenome-wide association studies (PheWAS) indicated limited associations with off-target traits, indicating a favorable safety profile for therapeutic intervention, with the exceptions of B4GALNT3, PM20D1, and CNN2. Integration of immune gene targets with pharmacological databases yielded three candidate compound, including NSC321521 (targeting HLA-DQA1), phenoxybenzamine (targeting GSTP1), and rimexolone (targeting BIN1). Among these compounds, Predicted blood-brain barrier permeability was observed only for phenoxybenzamine and rimexolone, with docking studies indicating stable interactions, such as those between NSC321521 and HLA-DQA1, phenoxybenzamine and GSTP1, and rimexolone and BIN1. This integrative approach highlights key immune&#x2011;cell&#x2011;specific genes involved in AD and proposes repurposable drugs with central nervous system potential, paving the way for more targeted immunomodulatory strategies in AD.

Humans↗

Estimating mutation parameters, population history and genealogy simultaneously from temporally spaced sequence data.

Molecular sequences obtained at different sampling times from populations of rapidly evolving pathogens and from ancient subfossil and fossil sources are increasingly available with modern sequencing technology. Here, we present a Bayesian statistical inference approach to the joint estimation of mutation rate and population size that incorporates the uncertainty in the genealogy of such temporally spaced sequences by using Markov chain Monte Carlo (MCMC) integration. The Kingman coalescent model is used to describe the time structure of the ancestral tree. We recover information about the unknown true ancestral coalescent tree, population size, and the overall mutation rate from temporally spaced data, that is, from nucleotide sequences gathered at different times, from different individuals, in an evolving haploid population. We briefly discuss the methodological implications and show what can be inferred, in various practically relevant states of prior knowledge. We develop extensions for exponentially growing population size and joint estimation of substitution model parameters. We illustrate some of the important features of this approach on a genealogy of HIV-1 envelope (env) partial sequences.

Decision Trees↗

Navigating Sampling Bias in Discrete Phylogeographic Analysis: Assessing the Performance of an Adjusted Bayes Factor.

Bayesian phylogeographic inference is widely used in molecular epidemiological studies to reconstruct the dispersal history of pathogens. Discrete phylogeographic analysis treats geographic locations as discrete traits and infers lineage transition events among them, and is typically followed by a Bayes factor (BF) test to assess the statistical support. In the standard BF (BFstd) test, the relative abundance of the involved trait states is not considered, which can be problematic in the case of unbalanced sampling. Existing methods to correct sampling bias in discrete phylogeographic analyses using continuous-time Markov chain (CTMC) model, often require additional epidemiological information to balance the sampling effort among locations. As such data is not necessarily available, alternative approaches that rely solely on available genomic data are needed. In this perspective, we assess the performance of a modification of the BFstd, the adjusted Bayes factor (BFadj), which incorporates information on the relative abundance of samples by location when inferring support for transition events and root location inference without requiring additional data. Using a simulation framework, we assess the statistical performance of BFstd and BFadj under varying levels of sampling bias, estimating their type I and type II error rates. Our results show that BFadj complements the BFstd by reducing type I errors at the cost increasing type II errors for inferred transition events, while improving type I and type II errors in root location inference. Our findings provide guidelines for implementing the complementary BFadj to detect and mitigate sampling bias in discrete phylogeographic inference using CTMC modeling.

Bayes Theorem↗

Generative models for discovering sparse distributed representations.

We describe a hierarchical, generative model that can be viewed as a nonlinear generalization of factor analysis and can be implemented in a neural network. The model uses bottom-up, top-down and lateral connections to perform Bayesian perceptual inference correctly. Once perceptual inference has been performed the connection strengths can be updated using a very simple learning rule that only requires locally available information. We demonstrate that the network learns to extract sparse, distributed, hierarchical representations.

Algorithms↗

Nonlinear statistical modeling and model discovery for cardiorespiratory data.

We present a Bayesian dynamical inference method for characterizing cardiorespiratory (CR) dynamics in humans by inverse modeling from blood pressure time-series data. The technique is applicable to a broad range of stochastic dynamical models and can be implemented without severe computational demands. A simple nonlinear dynamical model is found that describes a measured blood pressure time series in the primary frequency band of the CR dynamics. The accuracy of the method is investigated using model-generated data with parameters close to the parameters inferred in the experiment. The connection of the inferred model to a well-known beat-to-beat model of the baroreflex is discussed.

Algorithms↗

A proportional hazards model for incidence and induced remission of disease.

To assess the protective effects of a time-varying covariate, we develop a stochastic model based on tumor biology. The model assumes that individuals have a Poisson-distributed pool of initiated clones, which progress through predetectable, detectable mortal and detectable immortal stages. Time-independent covariates are incorporated through a log-linear model for the expected number of clones, resulting in a proportional hazards model for disease onset. By allowing time-dependent covariates to induce clone death, with rate dependent on a clone's state, the model is flexible enough to accommodate delayed disease onset and remission or cure of preexisting disease. Inference uses Bayesian methods via Markov chain Monte Carlo. Theoretical properties are derived, and the approach is illustrated through analysis of the effects of childbirth on uterine leiomyoma (fibroids).

Adult↗

Modeling the effects of a bidirectional latent predictor from multivariate questionnaire data.

Researchers often measure stress using questionnaire data on the occurrence of potentially stress-inducing life events and the strength of reaction to these events, characterized as negative or positive and assigned an ordinal ranking. In studying the health effects of stress, one needs to obtain measures of an individual's negative and positive stress levels to be used as predictors. Motivated by data of this type, we propose a latent variable model, which is characterized by event-specific negative and positive reaction scores. If the positive reaction score dominates the negative reaction score for an event, then the individual's reported response to that event will be positive, with an ordinal ranking determined by the value of the score. Measures of overall positive and negative stress can be obtained by summing the reactivity scores across the events that occur for an individual. By incorporating these measures as predictors in a regression model and fitting the stress and outcome models jointly using Bayesian methods, inferences can be conducted without the need to assume known weights for the different events. We propose an MCMC algorithm for posterior computation and apply the approach to study the effects of stress on preterm delivery.

Algorithms↗

Ancient differentiation of the H and I haplomes in diploid Hordeum species based on 5S rDNA.

5S rDNA clones from 12 South American diploid Hordeum species containing the HH genome and 3 Eurasian diploid Hordeum species containing the II genome, including the cultivated barley Hordeum vulgare, were sequenced and their sequence diversity was analyzed. The 374 sequenced clones were assigned to "unit classes", which were further assigned to haplomes. Each haplome contained 2 unit classes. The naming of the unit classes reflected the haplomes, viz. both the long H1 and short I1 unit classes were identified with II genome diploids, and both the long H2 and long Y2 unit classes were recognized in South American HH genome diploids. Based upon an alignment of all sequences or alignments of representative sequences, we tested several evolutionary models, and then subjected the parameters of the models to a series of maximum likelihood (ML) analyses and various tests, including the molecular clock, and to a Bayesian evolutionary inference analysis using Markov chain Monte Carlo (MCMC). The best fitting model of nucleotide substitution was the HKY+G (Hasegawa, Kishino, Yano 1985 model with the Gamma distribution rates of nucleotide substitutions). Results from both ML and MCMC imply that the long H1 and short I unit classes found in the II genome diploids diverged from each other at the same rate as the long H2 and long Y2 unit classes found in the HH genome diploids. The divergence among the unit classes, estimated to be circa 7 million years, suggests that the genus Hordeum may be a paleopolyploid.

DNA, Plant↗

Emergence of two novel HIV-1 Circulating Recombinant Forms (CRF190_0708 and CRF191_0708): molecular characterization and clinical insights from a five-year study in Yunnan, China.

BACKGROUND: To characterize HIV-1 molecular epidemiology and identify novel circulating recombinant forms (CRFs) among antiretroviral therapy (ART)-na&#xef;ve heterosexuals in Yunnan, China, and evaluate their clinical impact. METHODS: This study examined 636 HIV-1 pol sequences to analyze genetic diversity, pretreatment drug resistance (PDR), and transmission networks. Near full-length genomes were obtained to identify and characterize novel recombinants, with their evolutionary history inferred by Bayesian analysis. Co-receptor tropism was predicted, and the five-year clinical outcomes (including immune reconstitution and virologic response) of patients infected with the novel CRFs were compared. RESULTS: The most prevalent type identified was CRF08_BC, accounting for 50.16% of cases. The prevalence of drug resistance was 5.97% (38/636), with the K103N mutation being the most common. An analysis of transmission networks revealed that 52.2% (272/521) of clusters were associated with CRF07_BC and CRF08_BC. Two novel second-generation CRFs were identified: CRF190_0708, with an estimated time to the most recent common ancestor (tMRCA) of 1998.9, and CRF191_0708, with a more recent tMRCA ranging from 2009.5 to 2011.6. During the five-year follow-up period, viral rebound was observed in 7 patients in the CRF190_0708 group and in 1 patient in the CRF191_0708 group. Drug-resistance mutations (M184V and K103N) were detected in a subset of rebound cases in the CRF190_0708 group. CONCLUSIONS: This study identifies two novel HIV-1 recombinants, CRF190_0708 and CRF191_0708, highlighting ongoing viral evolution in Yunnan. Preliminary findings suggest possible clinical differences, warranting further investigation. Continued molecular surveillance is needed. TRIAL REGISTRATION: The clinical study was registered at ClinicalTrials.gov under the identifier NCT03852849. The date of registration was March 22, 2019.

Adult↗

Identification of a PRDM1-regulated T cell network to regulate atherosclerotic plaque inflammation.

BACKGROUND: Inflammation is a key driver of atherosclerosis, yet the mechanisms sustaining inflammation in human plaques remain poorly understood. This study uses a network-based approach to identify immune gene programs involved in the transition from low- to high-risk (rupture-prone) human atherosclerotic plaques. METHODS: Expression data from human carotid artery plaques, both stable (low-risk, n&#x2009;=&#x2009;16) and unstable (high-risk, n&#x2009;=&#x2009;27), were analyzed using Weighted Gene Co-expression Network Analysis (WGCNA). Bayesian network inference, operated on the eigengene values from the WGCNA, further extended the WGCNA analysis, and similarity to the signature of T cell subsets was validated in single-cell RNA sequencing data of human plaques, and a&#xa0;loss-of-function study in a mouse model of atherosclerosis. In silico drug repurposing was performed to identify potential therapeutic targets. RESULTS: Our analysis revealed a distinct gene module with a prominent T cell signature, particularly in unstable plaques. Key regulatory factors, RUNX3, IRF7 and in particular PRDM1, were significantly downregulated in plaque T cells from symptomatic versus asymptomatic patients, indicating a protective role. Additionally, as PRDM1 is downstream of IRF7, we opted for PRDM1 as a key target. T cell-specific Prdm1 deficiency in Western-type diet fed Ldlr knockout mice&#xa0;featured accelerated plaque progression. Finally, as PRDM1 targeting&#xa0;drugs are not yet available, we performed in silico drug repurposing, identifying EGFR inhibitors as promising therapeutic candidates. CONCLUSIONS: This study highlights a PRDM1-regulated T cell network that distinguishes high-risk from low-risk plaques and demonstrates the regulatory role of T cell PRDM1 in controlling atherosclerosis, positioning this pathway as a promising therapeutic target.

Plaque, Atherosclerotic↗

Phylogeographic epidemiology of Dabie bandavirus in East Asia: divergent transmission networks and genotype&#x2011;linked clinical severity.

BACKGROUND: Severe fever with thrombocytopenia syndrome (SFTS), caused by Dabie bandavirus (SFTSV), exhibits geographically decoupled incidence and fatality patterns across East Asia. We aimed to elucidate the distinct ecological drivers and phylogeographic dynamics underlying this inland-coastal epidemiological divergence. METHODS: Integrating 1820 high-quality global genomes of SFTSV with well-characterized clinical cohorts (936 patients) and nationwide surveillance data (27,457 cases) from China, we constructed a comprehensive analytical framework. Ecological modeling, Bayesian phylogeography, and genotype-phenotype association analyses were employed to trace the evolutionary trajectories and clinical implications of the virus. RESULTS: A pronounced "inland-high-incidence vs. coastal-high-fatality" pattern of SFTS was identified. The incidence of SFTS exhibited divergent sensitivities to meteorological factors; inland transmission was sensitive to thermal fluctuations, whereas coastal dynamics were constrained by a sunshine threshold (>&#x2009;200&#xa0;h/month). In contrast, spatial divergence in clinical severity correlated with the distribution of regional viral genetic structures. Inland regions mainly co-circulated genotypes A, C, and D, while coastal regions were dominated by genotype B. Zhejiang province was identified as a genetic hub with significantly higher recombination frequencies than inland regions (11.0% vs. 3.5%, P < 0.001). Bayesian phylogeographic inference indicated frequent lineage exchange of Zhejiang province in China with the Republic of Korea and Japan. Clinically, genotypes B and D were associated with elevated mortality in coastal and inland regions, respectively, suggesting that the severe coastal phenotype is shaped by its genotype B-dominated structure. Additionally, the RdRp-N828S mutation emerged as a robust molecular correlate of fatal outcomes, warranting further functional validation. CONCLUSIONS: Divergent meteorological factors and plausible maritime transmission networks may underlie the geographically decoupled epidemiology of SFTS. These findings highlight that risk assessment must extend beyond incidence alone and provide a phylogeographically informed framework for targeted surveillance and genotype-specific interventions in high-risk hotspots.

Humans↗

Population toxicokinetics of benzene.

In assessing the distribution and metabolism of toxic compounds in the body, measurements are not always feasible for ethical or technical reasons. Computer modeling offers a reasonable alternative, but the variability and complexity of biological systems pose unique challenges in model building and adjustment. Recent tools from population pharmacokinetics, Bayesian statistical inference, and physiological modeling can be brought together to solve these problems. As an example, we modeled the distribution and metabolism of benzene in humans. We derive statistical distributions for the parameters of a physiological model of benzene, on the basis of existing data. The model adequately fits both prior physiological information and experimental data. An estimate of the relationship between benzene exposure (up to 10 ppm) and fraction metabolized in the bone marrow is obtained and is shown to be linear for the subjects studied. Our median population estimate for the fraction of benzene metabolized, independent of exposure levels, is 52% (90% confidence interval, 47-67%). At levels approaching occupational inhalation exposure (continuous 1 ppm exposure), the estimated quantity metabolized in the bone marrow ranges from 2 to 40 mg/day.

Bayes Theorem↗