Search PubMedSearch

SEARCH · Search PubMed

Results for “Bayesian estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Genomic prediction and genome-wide association study for liver abscesses in crossbred beef cattle.

Liver abscesses are a concern in feedlot cattle, and little is known about the role of genetics in their development. This study aimed to estimate genetic parameters and to identify single-nucleotide polymorphisms (SNPs) associated with liver abscesses. Crossbred cattle representing 18 breeds in the U.S. Meat Animal Research Center Germplasm Evaluation Program were phenotyped for liver abscesses at slaughter (n&#x2005;=&#x2005;9,044). Seventeen percent of cattle had liver abscesses. These cattle had genotypes that were imputed to sequence variant genotypes. After filtering and quality control, 340,723 SNPs were used in the analysis. Liver abscess prevalence was modeled with a single-step genomic best linear unbiased prediction (ssGBLUP) threshold model using a Bayesian framework. The model included contemporary group (sex, treatment group, and slaughter date), additive genomic, and residual effects. Genomic heritability was 0.039 (95% highest posterior density&#x2005;=&#x2005;0.005, 0.081), which was very small. To assess prediction quality, a 5-fold random cross-validation structure was used. Method Linear Regression was used to assess accuracy, bias, and dispersion by comparing estimated breeding values (EBV) from full and reduced analyses. Cross-validation metrics showed EBV based on genotypes had 0.05 reliability (SD&#x2005;<&#x2005;0.01) with no bias relative to EBV based on genotypes and phenotypes. For the genome-wide association study, SNP effects were back calculated from the EBV solutions from ssGBLUP. No SNPs were associated with liver abscesses at a Benjamini-Hochberg adjusted 0.05 significance level. Although a large dataset was used, this result was because of the low genomic heritability and imprecise EBV used to calculate SNP effects. Based on these results, environmental factors contribute to most of the variation in liver abscesses. Genetic selection to reduce liver abscesses would be slow because of the low genomic heritability, measurement late in life, and inability to measure breeding animals. A faster approach would be finding additional environmental interventions that maintain animal performance.

Animals

Daily low-dose carboplatin or weekly carboplatin plus nab-paclitaxel for concurrent chemoradiotherapy in older patients with locally advanced non-small cell lung cancer (JCOG1914): A randomized phase 3 trial.

BACKGROUND: Daily low-dose carboplatin with concurrent thoracic radiotherapy is the standard treatment for older patients with unresectable locally advanced non-small cell lung cancer (LA-NSCLC) in Japan. METHODS: This open-label phase 3 trial was conducted at 38 institutions in Japan. Patients aged&#xa0;&#x2265;&#xa0;75&#xa0;years with LA-NSCLC were randomly assigned (1:1) to receive daily carboplatin (30&#xa0;mg/m2) or weekly carboplatin (area under the curve, 2&#xa0;mg&#xb7;min/mL) plus nab-paclitaxel (30&#xa0;mg/m2) with thoracic radiotherapy. Durvalumab maintenance therapy was recommended after treatment completion. The primary endpoint was overall survival, which was used to assess the non-inferiority of weekly carboplatin plus nab-paclitaxel compared to daily low-dose carboplatin. RESULTS: From December 2020 to March 2024, 124 patients were enrolled (carboplatin arm, 61 and carboplatin plus nab-paclitaxel arm, 63). In the planned interim analysis, the Bayesian predictive probability indicating the non-inferiority of carboplatin plus nab-paclitaxel compared with carboplatin in the final analysis was 8.0%, leading to early study termination for futility. The median overall survival was not estimable in the carboplatin arm; the estimated value in the carboplatin plus nab-paclitaxel arm was 26.1&#xa0;months (hazard ratio, 1.56; 95% confidence interval, 0.79-3.11; p&#xa0;=&#xa0;0.200). Two treatment-related and seven non-cancer-related deaths occurred in the carboplatin plus nab-paclitaxel arm. Patients in the carboplatin arm had better quality of life than those in the carboplatin plus nab-paclitaxel arm at 6&#xa0;weeks (odds ratio, 0.39; 95% confidence interval, 0.18-0.81; p&#xa0;=&#xa0;0.012). CONCLUSIONS: Daily low-dose carboplatin with concurrent thoracic radiotherapy remains the standard treatment for older patients with unresectable LA-NSCLC in Japan.

Humans

Twin azygotic test for the study of hereditary qualitative traits in twin populations.

Following previous formulations of a model of qualitative analysis of twin population data independent of zygosity, a new Bayesian approach has been developed. The present model can be applied to any qualitative genetic trait in twin population data, provided no specific source of variation be introduced by the twin condition, and allows not only estimation of the frequencies of mono- and dizygosity as well as the gene frequencies, but also verification of the trait's mode of inheritance.

Bayes Theorem

Multi-omics integration and colocalization analyses prioritize candidate molecular loci associated with hypothermia.

BACKGROUND: Hypothermia is a life-threatening condition lacking specific pharmacological treatments. This study aimed to prioritize genetically supported molecular loci associated with hypothermia and to explore their pharmacological tractability using multi-omics data. METHODS: Initially, 2532 druggable genes were curated from the Drug-Gene Interaction Database and established literature. These were cross-referenced with cis-eQTL and cis-pQTL datasets, encompassing 870,655 and 114,281 SNPs for blood, respectively, alongside 2379 shared SNPs across adipose, skeletal muscle, and heart tissues. Matched instrumental variables were integrated with hypothermia GWAS summary statistics for two-sample Mendelian randomization (MR) and Bayesian colocalization. Transcriptomic differential expression analysis (DEA) was subsequently conducted as an exploratory analysis of cold-exposure-associated expression changes. Database-derived compound annotations were systematically re-evaluated according to target specificity, established pharmacological mechanism, and concordance with the direction of the MR estimates. RESULTS: Among 671 gene-level MR tests, 36 genes reached nominal significance, whereas only ABCC8 remained significant after FDR correction. Colocalization was evaluable for 8 of these 36 genes, and 4 loci (COL18A1, SLC1A7, ADIPOQ, and MERTK) met the prespecified PP.H4>0.90 threshold. The remaining 28 loci were not evaluable because sufficient overlapping regional variants were unavailable after harmonization. Transcriptomic analysis identified altered expression of SLC1A3 and SLCO4A1 under cold exposure, although these findings did not directly validate the colocalization-supported loci. Re-evaluation of database-derived compound annotations did not identify any direct, selective, and directionally concordant drug-repurposing candidate for hypothermia. CONCLUSIONS: COL18A1, SLC1A7, ADIPOQ, and MERTK showed colocalization support among the 8 evaluable nominal MR-associated loci. Because colocalization coverage was limited, these genes should be regarded as preliminary candidate loci rather than established therapeutic targets. The pharmacological annotations were indirect, non-selective, unsupported, or directionally inconsistent and should be interpreted solely as hypothesis-generating information.

Bayesian colocalization

Bayesian classification of OXPHOS deficient skeletal myofibres.

Mitochondria are organelles in most human cells which release the energy required for cells to function. Oxidative phosphorylation (OXPHOS) is a key biochemical process within mitochondria required for energy production and requires a range of proteins and protein complexes. Mitochondria contain multiple copies of their own genome (mtDNA), which codes for some of the proteins and ribonucleic acids required for mitochondrial function and assembly. Pathology arises from genetic defects in mtDNA and can reduce cellular abundance of OXPHOS proteins, affecting mitochondrial function. Due to the continuous turn-over of mtDNA, pathology is random and neighbouring cells can possess different OXPHOS protein abundance. Estimating the proportion of cells where OXPHOS protein abundance is too low to maintain normal function is critical to understanding disease severity and predicting disease progression. Currently, one method to classify single cells as being OXPHOS deficient is prevalent in the literature. The method compares a patient's OXPHOS protein abundance to that of a small number of healthy control subjects. If the patient's cell displays an abundance which differs from the abundance of the controls then it is deemed deficient. However, due to the natural variation between subjects and the low number of control subjects typically available, this method is inflexible and often results in a large proportion of patient cells being misclassified. These misclassifications have significant consequences for the clinical interpretation of these data. We propose a single-cell classification method using a Bayesian hierarchical mixture model, which allows for inter-subject OXPHOS protein abundance variation. The model accurately classifies an example dataset of OXPHOS protein abundances in skeletal muscle fibres (myofibres). When comparing the proposed and existing model classifications to manual classifications performed by experts, the proposed model results in estimates of the proportion of deficient myofibres that are consistent with expert manual classifications.

Oxidative Phosphorylation

Genomic and clinical epidemiology of SARS-CoV-2 in coastal Kenya: insights into variant circulation, reinfection, and multiple lineage importations during a post-pandemic wave.

BACKGROUND: Between November 2023 and March 2024, coastal Kenya experienced another wave of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infections detected through our continued genomic surveillance. Herein, we report the clinical and genomic epidemiology of SARS-CoV-2 infections from 179 individuals (a total of 185 positive samples) residing in the Kilifi Health and Demographic Surveillance System (KHDSS) area (~&#x2009;900 km2). METHODS: We analyzed genetic, clinical, and epidemiological data from SARS-CoV-2 positive cases across pediatric inpatient, health facility outpatient, and homestead community surveillance platforms. Phylogenetic analyses were performed using maximum-likelihood and Bayesian frameworks. Temporal trends were summarized, comparisons conducted using Kruskal-Wallis and Wilcoxon tests, and associations examined using univariate and multivariable logistic regression models. RESULTS: Sixteen SARS-CoV-2 lineages within 3 subvariants [XBB.2.3-like (58.4%), JN.1-like (40.5%), and XBB.1-like (1.1%)] were identified. The symptomatic infection rate was estimated at 16.0% (95% CI, 11.1-23.9%) based on community testing regardless of symptom status and did not differ across the subvariants (p&#x2009;=&#x2009;0.13). The most common infection symptoms in community cases were cough (49.2%), fever (27.0%), sore throat (7.3%), headache (6.9%), and difficulty in breathing (5.5%). One case succumbed to the infection. Genomic analysis of the virus from serial positive samples confirmed repeat infections among 5 participants under follow-up (median interval 21&#xa0;days, range 16-95&#xa0;days); in 4 participants, the same virus lineage was responsible in both episodes, whereas 1 participant had a different lineage in the second compared with the first episode. Phylogenetic analysis including&#x2009;>&#x2009;18,000 contemporaneous global sequences provided evidence for at least 38 independent virus introduction events into the study area (KHDSS) during the wave, the majority likely originating in North America and Europe. CONCLUSIONS: Our study highlights that coastal Kenya, like most other localities, continues to face new SARS-CoV-2 infection waves characterized by circulation of new variants, multiple lineage importations, and reinfections. Locally, the virus may circulate unrecognized, as most infections are asymptomatic in part due to high population immunity after several waves of infection. Our findings highlight the need for sustained SARS-CoV-2 surveillance to inform appropriate public health responses, such as scheduled vaccination for populations at risk of severe infection.

COVID-19

Quantifying uncertainty of predictions from cancer progression models.

MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).

Humans

A flexible framework for robust and efficient Mendelian randomization with debiasing.

Mendelian randomization (MR) has been widely used to infer causal relationships between exposures and outcomes in epidemiological studies. However, classical MR assumptions can be violated when genetic variants are associated with outcomes through pathways other than the exposure, leading to uncorrelated and/or correlated pleiotropy. Additionally, measurement error arising from the inherent uncertainty in summary statistics obtained from large-scale genome-wide association studies can introduce bias into the causal effect estimate. To address these issues, we develop a debiased mixture inverse variance weighting ($\mathsf{dmIVW}$) method with three major advantages. First, it is capable of simultaneously handling various types of pleiotropy and eliminating the bias caused by uncertainty. Second, it can guard against distortion caused by invalid genetic variants while effectively harnessing their information. Third, our unified framework facilitates a fair comparison and combination of a series of submodels, encompassing several popular MR methods as special cases. Through real data applications, the effectiveness and robustness of $\mathsf{dmIVW}$ in estimating the causal effects of risk factors on common diseases are demonstrated.

Mendelian Randomization Analysis

Lineage-specific transmission and spatial clustering of Mycobacterium tuberculosis in Kaohsiung, Taiwan, in 2019-23: a population-based genomic study.

BACKGROUND: The epidemiology of tuberculosis in Taiwan has been influenced by the introduction of multiple Mycobacterium tuberculosis lineages and by the ageing of the population. We conducted a population-based study to investigate M tuberculosis transmission in Kaohsiung, a city in southern Taiwan. METHODS: In this study, we performed whole-genome sequencing (WGS) of M tuberculosis isolates from all culture-positive cases of tuberculosis notified in Kaohsiung between Jan 1, 2019 and Dec 31, 2023. We obtained routine epidemiological data for each case collected through the national tuberculosis control programme. We characterised the lineage composition of the isolate collection and evaluated genomic clustering of isolates, defined as a difference of 12 or fewer single-nucleotide polymorphisms. Univariable and multivariable logistic regression analyses were performed to estimate the odds of a case belonging to a genomic cluster based on host factors (age, sex, sputum smear status, and residential region) and pathogen factors (drug resistance status and strain lineage). Spatial aggregation of large genomic clusters (including greater than or equal to ten isolates) was assessed using a non-parametric statistical clustering method. We used a Bayesian transmission tree inference method to explore the patterns of age-dependent transmission. FINDINGS: During the study period, 5667 tuberculosis cases were notified in Kaohsiung, 4916 (86&#xb7;7%) of which were culture-positive. Of these 4916 cases, whole-genome sequencing was successfully performed for 4168 (84&#xb7;8%) isolates. 1219 (29&#xb7;2%) of 4168 individuals were female and 2947 (70&#xb7;7%) were male; the median age was 69&#xb7;7 years (IQR 57&#xb7;4-80&#xb7;7). The dominant lineages were lineage 1 (1749 [42&#xb7;0%] of 4168 isolates), lineage 2 (1510 [36&#xb7;2%]), and lineage 4 (905 [21&#xb7;7%]). 1069 (25&#xb7;6%) of 4168 were genomically linked and formed 287 clusters. Lineage 2 isolates had higher odds (aOR 2&#xb7;15 [95% CI 1&#xb7;80-2&#xb7;52]) than lineage 1 isolates of genomic clustering across all regions, whereas lineage 4 isolates had a significantly higher risk (2&#xb7;75 [1&#xb7;16-6&#xb7;89]) of genomic clustering than lineage 1 only in the rural northeast region, inhabited primarily by indigenous populations. Spatial clustering analysis corroborated these lineage-region interactions. Although younger adults (<35 years) had the highest individual-level odds (5&#xb7;64 [4&#xb7;16-7&#xb7;68]) of clustering in the logistic regression analysis compared with those aged 80 years or older, the transmission inference indicated that individuals aged 55-74 years were responsible for a greater proportion of inferred transmission events, contributing 50&#xb7;8% of all transmission events. INTERPRETATION: This sequencing study revealed that older adults (aged &#x2265;65 years) might have played a substantial and under-recognised role in the transmission of tuberculosis in Taiwan. The lineage-specific clustering and spatial patterns suggested that both pathogen characteristics and host demographics shaped tuberculosis transmission dynamics. These findings support the use of integrated genomic surveillance to guide precision tuberculosis control and motivate further research on age-specific transmission pathways and targeted interventions to advance tuberculosis elimination efforts. FUNDING: Taiwan National Health Research Institutes and Taiwan National Science and Technology Council.

Mycobacterium tuberculosis

IQ-NET: fast and accurate quartet phylogenetic inference using deep learning trained on empirical DNA alignments.

Phylogenetic inference is fundamental to modern biology, with many applications including evolutionary biology, epidemiology, and comparative genomics. While maximum likelihood and Bayesian methods remain the gold standard for phylogenetic analysis, they rely on simplifying assumptions and are computationally intensive. Recent machine learning approaches for phylogenetics offer speed advantages, but have several limitations: exclusive reliance on simulated data for training, inadequate handling of gaps, and sensitivity to input sequence order. Here, we introduce IQ-NET (Intelligent Quartet NETwork), a deep learning framework that solves these limitations to infer four-taxon trees. IQ-NET estimates both tree topology and branch lengths directly from gapped alignments. IQ-NET outperforms existing machine learning methods in terms of accuracy, and obtained a 24-fold speedup compared with the widely used maximum likelihood software, IQ-TREE. We finally introduce a pipeline using IQ-NET and the ASTRAL software to reconstruct a larger species tree, i.e., with more than four taxa.

Empirical data training

Comparative effectiveness of game-based learning modalities in nursing and medical education: a systematic review and Bayesian network meta-analysis.

BACKGROUND: Game-based learning (GBL) is increasingly used in healthcare education, but educators must choose among diverse modalities (e.g., quiz platforms, apps, serious games and metaverse environments). Comparative evidence on which modalities perform best across learning domains (knowledge, attitudes, and practice) remains limited. AIM: To compare the effects of distinct GBL modalities on knowledge, attitudes, and practice outcomes in nursing and medical education and to explore whether comparative effects differ by learner group (pre-licensure students and in-service professionals). DESIGN: PRISMA-NMA-aligned systematic review and Bayesian network meta-analysis. METHODS: We searched eight databases and trial registries through September 2, 2024, for randomized controlled trials comparing GBL with traditional teaching (TT). Outcomes were transformed to a 0-100 scale and analysed as change from baseline in Bayesian consistency models; random-effects models were selected using deviance information criterion (DIC). Risk of bias was assessed using RoB 2. We report mean differences (MDs) with 95% credible intervals (CrIs) versus TT, ranking probabilities, and subgroup NMAs by learner group. RESULTS: Thirty-one RCTs (n&#xa0;=&#xa0;3439) were included; 15 contributed complete data to the network. Risk of bias was low in 15 trials and raised some concerns in 16. The network was modest for knowledge (11 trials) and sparse for attitudes (3) and practice (4). Compared with TT, metaverse-based learning showed improved attitudes (MD 15; 95% CrI 12 to 18), based on a single trial. For knowledge and practice, Kahoot-based quizzes (MD 9.1; 95% CrI -8.9 to 27) and app-based learning (MD 4.6; 95% CrI -4.4 to 14) had the highest estimated mean improvements, but credible intervals were wide and included the null for most comparisons. Subgroup rankings differed by learner group, but several comparisons were imprecise and uncertainty was substantial, particularly in sparse networks. CONCLUSIONS: GBL modalities may improve learning outcomes compared with TT, but relative effects appear domain-specific and the certainty of rankings is limited by sparse evidence and imprecision. Future trials should prioritise head-to-head comparisons, robust outcome measurement, and longer-term retention and transfer outcomes in both student and in-service populations.

Humans

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics

Emergence of two novel HIV-1 Circulating Recombinant Forms (CRF190_0708 and CRF191_0708): molecular characterization and clinical insights from a five-year study in Yunnan, China.

BACKGROUND: To characterize HIV-1 molecular epidemiology and identify novel circulating recombinant forms (CRFs) among antiretroviral therapy (ART)-na&#xef;ve heterosexuals in Yunnan, China, and evaluate their clinical impact. METHODS: This study examined 636 HIV-1 pol sequences to analyze genetic diversity, pretreatment drug resistance (PDR), and transmission networks. Near full-length genomes were obtained to identify and characterize novel recombinants, with their evolutionary history inferred by Bayesian analysis. Co-receptor tropism was predicted, and the five-year clinical outcomes (including immune reconstitution and virologic response) of patients infected with the novel CRFs were compared. RESULTS: The most prevalent type identified was CRF08_BC, accounting for 50.16% of cases. The prevalence of drug resistance was 5.97% (38/636), with the K103N mutation being the most common. An analysis of transmission networks revealed that 52.2% (272/521) of clusters were associated with CRF07_BC and CRF08_BC. Two novel second-generation CRFs were identified: CRF190_0708, with an estimated time to the most recent common ancestor (tMRCA) of 1998.9, and CRF191_0708, with a more recent tMRCA ranging from 2009.5 to 2011.6. During the five-year follow-up period, viral rebound was observed in 7 patients in the CRF190_0708 group and in 1 patient in the CRF191_0708 group. Drug-resistance mutations (M184V and K103N) were detected in a subset of rebound cases in the CRF190_0708 group. CONCLUSIONS: This study identifies two novel HIV-1 recombinants, CRF190_0708 and CRF191_0708, highlighting ongoing viral evolution in Yunnan. Preliminary findings suggest possible clinical differences, warranting further investigation. Continued molecular surveillance is needed. TRIAL REGISTRATION: The clinical study was registered at ClinicalTrials.gov under the identifier NCT03852849. The date of registration was March 22, 2019.

Adult

Assessment of Genetic Diversity and Population Structure on Azadirachta indica A. Juss. in an Urban Metropolitan: Ahmedabad, India.

Azadirachta indica (A. indica) A. Juss., commonly known as Neem, is a valuable multipurpose tree with profound medicinal properties and socioeconomic importance, widely recognized since ancient Ayurvedic times. Despite its prominence, knowledge about its genetic diversity within the metropolitan area of Ahmedabad is limited. This study marks the first in-depth exploration of the genetic diversity and population structure of A. indica in Ahmedabad. The authenticity of the species was validated through DNA barcoding, and a Geographical Information System (GIS) was used to collect the samples. A total of 35 A. indica accessions were analyzed using five Inter Simple Sequence Repeat (ISSR) primers. Genetic diversity and population structure were evaluated using Inter Simple Sequence Repeat (ISSR) markers through polymorphism assessment, clustering, ordination, and Bayesian population structure analyses. ISSRs revealed a high level of polymorphism (75.66%), indicating substantial genetic variability among accessions. An analysis of genetic diversity indices revealed low to moderate diversity (Hs&#x2009;=&#x2009;0.14, Ht&#x2009;=&#x2009;0.217, I&#x2009;=&#x2009;0.217). Analysis of Molecular Variance (AMOVA) analysis depicted 81% variation within the population and 19% among the population. Low to moderate genetic differentiation (Gst&#x2009;=&#x2009;0.319) and moderate gene flow (Nm&#x2009;=&#x2009;1.06) indicated that urban development has not hindered gene flow among populations. Mantel's test revealed a weak but significant correlation between genetic and geographic distances, suggesting limited isolation by distance. The estimated &#x394;K using STRUCTURE exhibited two subpopulations, representing two gene pools for A. indica accessions (K&#x2009;=&#x2009;2). Collectively, these patterns indicate that urbanization has not severely disrupted genetic connectivity in A. indica, reflecting its resilience and adaptive potential in a metropolitan environment. These findings provide pivotal knowledge for further understanding the genetic diversity and population structure of A. indica in one of the fastest-growing cities in India, which can be utilized for new breeding programmes, sustainable development and future conservation strategies around the globe.

India

On the deconvolution of exponential response functions.

The deconvolution or unfolding of exponential response functions from experimental data has been examined through the use of a Bayesian based algorithm. The algorithm, which is founded upon the concepts of probability, ensures positivity of solution. This constraint leads to a significant reduction in the growth of statistical noise in deconvolved data when compared with the more common linear unfolding techniques. The algorithm is an iterative procedure which, in the absence of statistical noise, can ultimately result in complete signal recovery. When noise is present one must balance the degree with which the response function is removed against the growth in the noise and, at some point, terminate the iterative process. Criteria for determining the point at which this 'best estimate' is attained are examined and an operationally realisable test is given. Comparison of results is made with the inverse filter solution which, for an exponential response function, is shown to consist of the sum of the observed data and its first derivative.

Mathematics

ScITree: Scalable Bayesian inference of transmission tree from epidemiological and genomic data.

Phylodynamic models capture joint epidemiological-evolutionary dynamics during an outbreak, providing a powerful tool to enhance understanding and management of disease transmission. Existing phylodynamic approaches, however, mostly rely on various non-mechanistic or semi-mechanistic approximations of the underlying epidemiological-evolutionary process. Previous work by Lau and colleagues has shown that full Bayesian mechanistic models, without relying on these approximations, can enable highly accurate joint inference of the epidemiological-evolutionary dynamics including the unobserved transmission tree. However, the Lau method faces major computational bottlenecks. As the volume of genomic data collected during outbreaks continues to grow, it is crucial to develop scalable yet accurate phylodynamic methods. Here we propose a new Bayesian phylodynamic model, overcoming the major scalability issue in the previous method and enabling a readily deployable, yet accurate, phylodynamic modeling framework. Specifically, we develop a scalable spatio-temporal phylodynamic framework for inferring the transmission tree (ScITree) and other key epidemiological parameters considering the infinite sites assumption in modeling mutation on the sequence level, in contrast to the Lau method in which mutation was modeled explicitly on the nucleotide level. Our approach features full Bayesian implementation utilizing an exact likelihood to mechanistically integrate epidemiological and evolutionary processes. We develop a computationally-efficient data-augmentation Markov Chain Monte Carlo algorithm, inferring key model parameters and unobserved dynamics including the transmission tree. We assess performance of our method using multiple simulated outbreak datasets. Our results indicate that our method can achieve high inference accuracy, comparable to the performance of the Lau method. Additionally, our method scales significantly more efficiently for large outbreaks, with computing time increasing linearly with outbreak size, compared to the exponential scaling of the Lau method. We also demonstrate our method's utility by applying our validated modeling framework to a dataset describing a foot-and-mouth disease outbreak in the UK. Our results show that our method is able to generate estimates of the transmission dynamics consistent with those from the prior method, further demonstrating the robustness of our new approach. In summary, our method provides a computationally-efficient, highly scalable, accurate modeling framework for inferring the joint spatio-temporal dynamics of epidemiological and evolutionary processes, facilitating timely and effective outbreak responses in space and time. Our method is implemented in our R package ScITree.

Bayes Theorem

Comparison on Major Gene Mutations Related to Rifampicin and Isoniazid Resistance between Beijing and Non-Beijing Strains of Mycobacterium tuberculosis: A Systematic Review and Bayesian Meta-Analysis.

Objective: The Beijing strain of Mycobacterium tuberculosis (MTB) is controversially presented as the predominant genotype and is more drug resistant to rifampicin and isoniazid compared to the non-Beijing strain. We aimed to compare the major gene mutations related to rifampicin and isoniazid drug resistance between Beijing and non-Beijing genotypes, and to extract the best evidence using the evidence-based methods for improving the service of TB control programs based on genetics of MTB. Method: Literature was searched in Google Scholar, PubMed and CNKI Database. Data analysis was conducted in R software. The conventional and Bayesian random-effects models were employed for meta-analysis, combining the examinations of publication bias and sensitivity. Results: Of the 8785 strains in the pooled studies, 5225 were identified as Beijing strains and 3560 as non-Beijing strains. The maximum and minimum strain sizes were 876 and 55, respectively. The mutations prevalence of rpoB, katG, inhA and oxyR-ahpC in Beijing strains was 52.40% (2738/5225), 57.88% (2781/4805), 12.75% (454/3562) and 6.26% (108/1724), respectively, and that in non-Beijing strains was 26.12% (930/3560), 28.65% (834/2911), 10.67% (157/1472) and 7.21% (33/458), separately. The pooled posterior value of OR for the mutations of rpoB was 2.72 ((95% confidence interval (CI): 1.90, 3.94) times higher in Beijing than in non-Beijing strains. That value for katG was 3.22 (95% CI: 2.12, 4.90) times. The estimate for inhA was 1.41 (95% CI: 0.97, 2.08) times higher in the non-Beijing than in Beijing strains. That for oxyR-ahpC was 1.46 (95% CI: 0.87, 2.48) times. The principal patterns of the variants for the mutations of the four genes were rpoB S531L, katG S315T, inhA-15C > T and oxyR-ahpC intergenic region. Conclusion: The mutations in rpoB and katG genes in Beijing are significantly more common than that in non-Beijing strains of MTB. We do not have sufficient evidence to support that the prevalence of mutations of inhA and oxyR-ahpC is higher in non-Beijing than in Beijing strains, which provides a reference basis for clinical medication selection.

Isoniazid

Navigating Sampling Bias in Discrete Phylogeographic Analysis: Assessing the Performance of an Adjusted Bayes Factor.

Bayesian phylogeographic inference is widely used in molecular epidemiological studies to reconstruct the dispersal history of pathogens. Discrete phylogeographic analysis treats geographic locations as discrete traits and infers lineage transition events among them, and is typically followed by a Bayes factor (BF) test to assess the statistical support. In the standard BF (BFstd) test, the relative abundance of the involved trait states is not considered, which can be problematic in the case of unbalanced sampling. Existing methods to correct sampling bias in discrete phylogeographic analyses using continuous-time Markov chain (CTMC) model, often require additional epidemiological information to balance the sampling effort among locations. As such data is not necessarily available, alternative approaches that rely solely on available genomic data are needed. In this perspective, we assess the performance of a modification of the BFstd, the adjusted Bayes factor (BFadj), which incorporates information on the relative abundance of samples by location when inferring support for transition events and root location inference without requiring additional data. Using a simulation framework, we assess the statistical performance of BFstd and BFadj under varying levels of sampling bias, estimating their type I and type II error rates. Our results show that BFadj complements the BFstd by reducing type I errors at the cost increasing type II errors for inferred transition events, while improving type I and type II errors in root location inference. Our findings provide guidelines for implementing the complementary BFadj to detect and mitigate sampling bias in discrete phylogeographic inference using CTMC modeling.

Bayes Theorem