Search PubMedSearch

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

A Bayesian group sequential design for a multiple arm randomized clinical trial.

Group sequential designs for randomized clinical trials allow analyses of accruing data. Most group sequential designs in the literature concern the comparison of two treatments and maintain an overall prespecified type I error. As the number of treatments increases, however, so does the probability of falsely rejecting the null hypothesis. Bayesian statisticians concern themselves with the observed data and abide by the likelihood principle. As long as previous analyses do not change the likelihood, these analyses do not change Bayesian inference. In this paper, we discuss a group sequential design for a proposed randomized clinical trial comparing four treatment regimens. Bayesian ideas underlie the design and posterior probability calculations determine the criteria for stopping accrual to one or more of the treatments. We use computer simulation to estimate the frequentists properties of the design, information of interest to many of our collaborators. We show that relatively simple posterior probability calculations, along with simulations to calculate power under alternative hypotheses, can produce appealing designs for randomized clinical trials.

Bayes Theorem

The diagnosis of polyarteritis nodosa. I. A literature-based decision analysis approach.

We investigated diagnostic testing in polyarteritis nodosa (PAN) by calculating, from published data, the sensitivity and specificity of visceral angiography and muscle, nerve, testicle, kidney, and liver biopsy. Test sequence strategies were constructed by Bayesian inference using a computer program written for this purpose. Test sequences were compared with an aggressive strategy consisting of repeated tests until there was a positive finding or until the available tests were exhausted, and a conservative strategy consisting of 1 biopsy procedure plus angiography. The Bayesian analysis agreed most closely with the conservative approach for most prior probabilities (degree of suspicion) that a patient had PAN. The aggressive strategy had an overall sensitivity of 90% and specificity of 91%, whereas the conservative strategy was 85% sensitive and 96% specific. Furthermore, the aggressive strategy was more costly ($2,986 versus $1,961) and had a higher rate of morbidity (3.8 versus 2.7 days of hospitalization per patient evaluated) than did the conservative strategy. The mortality rates of both strategies were equivalent (approximately 0.05 deaths per hundred patients evaluated). The per-case cost of diagnosis increased as prevalence decreased, and at 10% prevalence, the aggressive strategy cost more than $17,000 per case diagnosed. Sensitivity analysis revealed that the strategies were moderately affected by the test characteristics, within reasonable assumptions, but that the differences in conservative and aggressive approaches remained. Thus, our analysis based on available data and the assumption of test independence suggests that the preferred diagnostic evaluation of patients with symptoms suggestive of PAN consists, in most cases, of a single biopsy procedure, with angiographic evaluation if necessary.

Costs and Cost Analysis

Hypermedia and randomized algorithms for medical expert systems.

KNET is an environment for constructing probabilistic, knowledge-intensive systems within the axiomatic framework of decision theory. The KNET architecture defines a complete separation between the hypermedia user interface on the one hand, and the representation and management of expert opinion on the other. KNET offers a choice of algorithms for probabilistic inference. We and our coworkers have used KNET to build consultation systems for lymph-node pathology, bone-marrow transplantation therapy, clinical epidemiology, and alarm management in the intensive-care unit. Most important, KNET contains a randomized approximation scheme (RAS) for the difficult and almost certainly intractable problem of Bayesian inference. Our algorithm can, in many circumstances, perform efficient approximate inference in large and richly interconnected models of medical diagnosis. In this article, we describe the architecture of KNET, construct a randomized algorithm for probabilistic inference, and analyze the algorithm's performance. Finally, we characterize our algorithms' empiric behavior and explore its potential for parallel speedups. From design to implementation, then, KNET demonstrates the crucial interaction between theoretical computer science and medical informatics.

Algorithms

The problem of multiple inference in studies designed to generate hypotheses.

Epidemiologic research often involves the simultaneous assessment of associations between many risk factors and several disease outcomes. In such situations, often designed to generate hypotheses, multiple univariate hypothesis-testing is not an appropriate basis for inference. The number of true positive associations in a collection of many associations can be estimated by comparing the observed distribution of p values for the positive associations to a theoretical uniform distribution, or to the observed distribution of negative associations, or to an empiric randomization distribution. None of these approaches, however, will distinguish the true from the false positive associations. Various criteria for selecting a subset of associations to report are considered by the authors, including Bonferoni adjustment of p values, splitting the sample for searching and testing, Bayesian inference, and decision theory. The authors prefer an approach in which all associations in the data are reported, whether significant or not, followed by a ranking in order of priority for investigation using empirical Bayes techniques. Methods are illustrated by application to preliminary data from a study aimed at identifying hitherto unsuspected occupational carcinogens.

Bayes Theorem

The contributions of Jerome Cornfield to the theory of statistics.

This paper is a review of the contributions of Jerome Cornfield to the theory of statistics. It discusses several highlights of his theoretical work as well as describing his philosophy relating theory to application. The three areas discussed are: linear programming, urn sampling and its generalizations to the analysis of variance, and Bayesian inference. It is not widely known that Jerome Cornfield was perhaps the first to formulate and approximately solve the linear programming problem in 1941. His formulation was made for the famous "Diet Problem". An early publication introduced the method of indicator random variables in the context of urn sampling. This simple method allowed straightforward calculations of the low order moments for estimates arising from sampling finite populations and was later generalized to the two-way analysis of variance. The application of the urn sampling model to the analysis of variance served to illuminate how one chooses proper error terms for making tests in the analysis of variance table. Jerome Cornfield's philosophy on applications of statistics was dominated by a Bayesian outlook. His theoretical contributions in the past two decades were mainly concerned with the development of Bayesian ideas and methods. A brief survey is made of his main contributions to this area. A particularly noteworthy result was his demonstration that for the two-sample slippage problem of location, the likelihood function under a permutation setting is uninformative for the slippage parameter. However, the posterior distribution differs from the prior distribution despite the fact that the likelihood is uninformative.

Bayes Theorem

Plastome evolution and phylogenomic relationships in Ajuga (Lamiaceae, Ajugoideae).

BACKGROUND: Ajuga is currently known to include approximately 69 species, with a combined distribution extending throughout Eurasia, Africa, and Australia. Its popularity and significance are largely based on an extensive history of medicinal and horticultural use. It is divided into two sections based on morphological characters, and this sectional classification is also reflected in pronounced geographic patterns. Although previous studies have largely focused on Ajuga sect. Ajuga in East Asia, A. sect. Chamaepithys, which ranges from the Mediterranean to Central Asia, remains insufficiently sampled, thereby limiting a comprehensive understanding of infrageneric sectional relationships within the genus. Here, we generated complete plastid genomes for 12 species representing both sections of the genus and used these data to characterize plastome structure and infer evolutionary relationships. RESULTS: In this study, 21 Ajuga plastomes were analyzed, including 12 newly sequenced plastomes and 9 previously published plastomes representing 19 species. Comparative analyses showed that all plastomes exhibited a highly conserved quadripartite structure, with genome sizes ranging from 149,963 to 150,740 bp and GC contents varying from 38.2% to 38.3%. Each plastome contained 133 genes, including 88 protein-coding genes, 37 transfer RNA genes, and 8 ribosomal RNA genes. The boundaries between the inverted repeat (IR) and single-copy (SC) regions were also highly conserved across species. In addition, 796 simple sequence repeats (SSRs), 874 long repeat sequences (LRSs), and 12 highly variable regions (ccsA-ndhD, ndhF-rpl32, petA-psbJ, rpl32-trnL-UAG, rps2-rpoC2, trnH-GUG-psbA, trnK-UUU-rps16, trnP-UGG-psaJ, trnT-UGU-trnL-UAA, ycf15-trnL-CAA, ndhF, and ycf1) were identified among the 21 plastomes. Phylogenetic analyses based on four datasets and conducted using Maximum Likelihood and Bayesian Inference recovered two major clades corresponding to the traditionally recognized sectional classification, with one distributed from the Mediterranean to Central Asia and the other in East Asia. CONCLUSION: This study represents the most comprehensive plastome-based sampling of Ajuga to date, including representative species from the Mediterranean, Central Asia, and East Asia. Our results have significantly enhanced our understanding of its infrageneric relationships. The plastome resources generated in this study provide a valuable foundation for future research on species delimitation, phylogeny, and the evolutionary history of Ajuga.

Phylogeny

Insights Into the Structural Features, Codon Usage Patterns, and Phylogenetic Analysis in Neoniphon argenteus (Teleostei: Holocentriformes) Based on Complete Mitochondrial Genome.

Neoniphon argenteus, a widely distributed nocturnal coral reef fish in the family Holocentridae, plays an important role in maintaining coral reef ecosystem health, yet its phylogenetic position remains poorly resolved. To bridge this gap, we sequenced and analyzed the complete mitochondrial genome of a specimen from the South China Sea to characterize its structural features, codon usage patterns, and phylogenetic relationships. The 16,569 bp mitogenome (GenBank: PP190474.1) encodes 13 protein-coding genes (PCGs), 22 tRNAs, two rRNAs, and two non-coding regions, exhibiting a distinct A + T bias. All tRNAs fold into typical cloverleaf secondary structures except tRNA-Ser (AGN), which lacks the dihydrouridine (DHU) arm. The control region contains palindromic motifs (TACAT/ATGTA) capable of forming hairpin structures and five conserved sequence blocks, whereas the OL region harbors a conserved 5'-GCCGG-3' motif. RSCU analysis revealed 31 frequently used codons (RSCU > 1) with a pronounced preference for A/C-ending codons. The ΔRSCU method identified 10 candidate optimal codons (GCA, CAA, GAA, GGA, AUU, CUA, CCA, CGA, ACA, and GUC). Selection pressure analysis using EasyCodeML and site-specific models indicated that all PCGs are predominantly under purifying selection, with no significant evidence of pervasive positive selection. ND6 exhibited elevated pairwise Ka/Ks ratios (mean = 1.209 ± 0.047), consistent with reduced selective constraint rather than adaptive evolution. Phylogenetic analysis of 19 Holocentriformes species using maximum likelihood and Bayesian inference with partitioned models based on 13 PCGs and two rRNA genes (12S and 16S) assigned all taxa to two well-supported subfamilies (Holocentrinae and Myripristinae). Within Holocentrinae, Neoniphon species form a monophyletic clade nested within a paraphyletic Sargocentron, suggesting that the genus Sargocentron as currently defined is not monophyletic. This study provides useful baseline molecular data for further exploration of the evolutionary history of N. argenteus and other members of Holocentriformes.

Holocentridae

Comparison of the information in two lung function experiments.

The amount of ventilation relative to perfusion (the ventilation-perfusion ratio) received by the lung is a useful indicator of the efficiency of lung function. Two alternative techniques for recovering the ventilation-perfusion ratio are outlined. While both techniques rely on the use of inert gases, one is well established and the other is only in a developmental stage. This paper focuses on a comparison of the amount of statistical information provided by these two techniques about the ventilation-perfusion ratio. The criterion applied here for measuring amount of information has roots in communication theory and uses ideas inherent to Bayesian inference.

Bayes Theorem

Probability and the patient state space.

This paper describes work to develop a model-based system to support clinical decision-making. In previous articles, we have developed (from 695 measurement sets obtained from 148 patients) a physiologic state classification based on a set of 11 cardiovascular and metabolic measurements. There is an R or reference state, for stable ICU patients. Patients under (operative, traumatic, or compensated septic) stress, or with (septic or hepatic) metabolic, respiratory, or cardiac insufficiency are in the A, B, C, or D states, respectively. We wished to make the state easier to measure and eventually available continuously, automatically, and noninvasively, as well as reflecting a wider group of bodily systems. The 5 centers define a 4 dimensional affine subspace, designated the cardiovascular state space. Using eigenvector analysis, we have found four new derived physiologic variables CV1, CV2, CV3, and CV4 that span the state space. We have fit sets of linear regression equations that allow the patient's position in the state space, and therefore his state, to be determined from more easily obtainable sets of measurements. Further, we selected 1966 measurement sets from 512 patients at two hospitals. We used the data from 250 of these patients to define 13 prototypical types, namely survivors and deaths from various combinations of sepsis, cardiogenic decompensation, cirrhosis, and pneumonitis, following trauma or general surgery. For any future patient, the statistical theory of Bayesian inference allows one to infer back from the measurements observed to the probability of his being of any of these types and of surviving or dying. We used this method to predict the outcome of the other 262 patients, prospectively. Statistically, the predictions of survival or death were not significantly different from the actual. For individual patients, the method predicts a clinical course that closely follows the actual episodes in their history. These results confirm and explain the validity of the concept of the patient state and make the state easier to compute. The patient state and the probability plot together help to stage, select, and evaluate therapy. They do not replace the clinician's judgement, but rather are tools that help the clinician to exercise judgement.

Adult

Medical expert systems based on causal probabilistic networks.

Causal probabilistic networks (CPNs) offer new methods by which you can build medical expert systems that can handle all types of medical reasoning within a uniform conceptual framework. Based on the experience from a commercially available system and a couple of large prototype systems, it appears that CPNs are now an attractive alternative to other methods. A CPN is an intensional model of a domain, and it is therefore conceptually much closer to qualitative reasoning systems and to simulation systems than to rule-based or logic-based systems. Recent progress in Bayesian inference in networks has yielded computationally efficient methods. The inference method used follows the fundamental axioms of probability theory, and gives a sound framework for causal and diagnostic (deductive and abductive) reasoning under uncertainty. Experience with the prototypes indicates that it may be possible to use decision theory as a rational approach to test planning and therapy planning. The way in which knowledge is acquired and represented in CPNs makes it easy to express 'deep knowledge' for example in the form of physiological models, and the facilities for learning make it possible to make a smooth transition from expert opinion to statistics based on empirical data.

Artificial Intelligence

Mitochondrial genome characteristics and phylogenetic analysis of Ramaria longispora.

This study, for the first time, assembled and annotated the complete mitochondrial genome of R. longispora using high-throughput sequencing technology. The genome is a circular molecule with a total length of 157,712 bp and a GC content of 31.55%. It encodes 71 genes, including 15 core protein-coding genes (PCGs), 25 transfer RNA (tRNA) genes, 2 ribosomal RNA (rRNA) genes, 5 free-stranding open reading frames (ORFs), and 24 intronic ORFs. Among these, most free-stranding ORFs have unknown functions but include a DNA polymerase gene, while the intronic ORFs primarily encode LAGLIDADG and GIY-YIG endonucleases. The mitochondrial genome contains 39 introns. Phylogenetic analyses based on 15 core PCGs using Bayesian inference (BI) and maximum likelihood (ML) methods revealed that this R. longispora is most closely related to Ramaria flavescens and Ramaria ichnusensis. This study provides foundational data for mitochondrial genome research in the Ramaria genus and offers important references for taxonomic and evolutionary studies of this group.

Mitochondrial genome

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs

Detecting Interspecific Positive Selection Using Convolutional Neural Networks.

Traditional statistical methods using maximum likelihood and Bayesian inference can detect positive selection from an interspecific phylogeny and a codon sequence alignment based on model assumptions, but they are prone to false positives due to alignment errors and can lack power. These problems are particularly pronounced when faced with high levels of indels and divergence. To address these issues, we trained and tested convolutional neural network models on simulated data and achieved higher accuracy in detecting selection across a specific range of phylogenetic scenarios and evolutionary modes. This advantage is particularly evident when performing inference on noisy data prone to misalignments. Our method shows some ability to account for these errors, where most statistical frameworks fail to do so in a tractable manner. We explore the generalizability of our convolutional neural network models to unseen evolutionary scenarios and identify future avenues to achieve broader utility. Once trained, our convolutional neural network model is faster at test time, making it a scalable alternative to traditional statistical methods for large-scale, multigene analyses. In addition to binary classification (inference of the presence or absence of positive selection during the evolution of the sequences), we use saliency maps to understand what the model learns and observe how this could be leveraged for sitewise inference of positive selection.

Neural Networks, Computer

Mapping Sub-National Respiratory Virus Circulation in Cambodia Using Metatranscriptomic Sequencing: A Multi-Center Hospital-Based Surveillance Study.

BACKGROUND: Genomic surveillance can guide early detection of and response to emerging epidemics. Metatranscriptomic sequencing was used to investigate sub-national respiratory virus circulation in Cambodia from 2020 to 2023. METHODS: Nasopharyngeal swabs were collected from individuals aged 2 months to 65 years with influenza-like illness in four Cambodian hospitals. Metatranscriptomic data were generated by short-read RNA sequencing. Bernoulli space-time scan statistics were used to identify temporal virus clusters. Bayesian inference of phylogenetic trees was used to compute divergence times for temporally clustered, highly represented viruses (influenza A/H3N2 and B, Betacoronavirus 1, respiratory syncytial virus [RSV] A and B), and publicly available global influenza virus genomes. RESULTS: Of 1093 individuals, 499 (45.7%) had detectable respiratory viruses belonging to 68 distinct species. Moderate (N > 20) discrete time-clusters were noted of RSV-A (37 cases), Betacoronavirus 1 (21 cases), RSV-B (22 cases), and A/H3N2 (30 cases). The posterior median of time to most recent common ancestor ranged from 0.71 years (95% HPD 0.38-1.10) for Betacoronavirus 1 and 1.31 years (95% HPD 0.60-3.20) for A/H3N2, to 2.75 years (1.82-4.26) for RSV-A and 4.79 years (2.39-7.74) for RSV-B. A/H3N2 and influenza B virus genomes mapped to clades 3C.2a1b.2a.2a and Victoria 1A.3a.2, respectively, and inter-mixed with concurrent global strains. CONCLUSIONS: Multiple respiratory viruses circulated at a sub-national level in Cambodia from 2020 to 2023 despite pandemic disruptions. Influenza virus population diversity decreased during the height of lockdown but recovered in mid-2022. Re-emerging influenza strains were distinct from historically circulating strains and clustered with contemporaneous global variants, suggesting multiple external introductions.

Humans

On avoiding statistical bias in linkage-based counselling.

Using the Succession Rule of Laplace (1795) and related reasoning, this paper shows how to give unbiased counselling to patients when predictions are to be based on small samples. The recombination fraction can be regarded as a probability parameter, theta, which itself has a probability distribution between the limits of 0 and 1/2. The probability of a recombinant, P(Rec), is not numerically equal to the maximum likelihood estimate of theta, nor is it numerically equal to the maximum posterior probability estimate in Bayesian inference. Rather it is equal to the infinite sum of all possible theta values, each weighted according to its probability density p(theta) which denotes the relative probability that that theta value is the true one. The various published proposals for obtaining an unbiased estimate of theta are shown to be equivalent one to another, except for the simplifying approximations used.

Bias

Contrasting Patterns of Connectivity Between Populations of Euphotic and Mesophotic Hydroids in Reunion Island Support the Deep Reef Refuge Hypothesis.

In the context of coral reef decline, mesophotic coral ecosystems (MCEs, 30-150 m) offer hope for the recovery of degraded euphotic reefs. The Deep Reef Refuge Hypothesis (DRRH) postulates the potential of mesophotic reefs to reseed euphotic reefs. This hypothesis needs to be further tested by estimating connectivity along the depth gradient. Mesophotic data are lacking worldwide, particularly in the southwestern Indian Ocean (SWIO). Here, using a total of 2218 samples collected at depths ranging from 10 to 103 m, we estimated the connectivity of 7 hydroid species sampled at euphotic, upper, and lower mesophotic depths around Reunion Island using a multi-species comparative framework. Population genetic analyses using 8-17 microsatellite markers per species (80 markers in total) as well as Bayesian inference were performed to estimate population structure and contemporary migration rates to highlight connectivity patterns and directionality of gene flow between depths. The results revealed three main genetic patterns depending on the species: a horizontal stepping stone pattern between areas around the island, a vertical stepping stone pattern between adjacent depths, and a quasi-panmictic pattern. Each species showed some specificity within these patterns, but overall, at least 4 of the 7 species support the assumption of vertical connectivity from the Deep Reef Refuge Hypothesis, highlighting the importance of studying multiple species. The existence of vertical connectivity between euphotic and mesophotic depths in the southwestern Indian Ocean confirms the importance of mesophotic coral ecosystems for conservation efforts and our global understanding of coral reef ecosystem dynamics.

Animals

Experiencing and perceiving visual surfaces.

A theoretical framework is proposed to understand binocular visual surface perception based on the idea of a mobile observer sampling images from random vantage points in space. Application of the generic sampling principle indicates that the visual system acts as if it were viewing surface layouts from generic not accidental vantage points. Through the observer's experience of optical sampling, which can be characterized geometrically, the visual system makes associative connections between images and surfaces, passively internalizing the conditional probabilities of image sampling from surfaces. This in turn enables the visual system to determine which surface a given image most strongly indicates. Thus, visual surface perception can be considered as inverse ecological optics based on learning through ecological optics. As such, it is formally equivalent to a degenerate form of Bayesian inference where prior probabilities are neglected.

Depth Perception

Effect of atypical antibiotic resistance on microorganism identification by pattern recognition.

We classified microorganisms from the clinical laboratory by using information provided by the Gram stain and antibiotic sensitivity profiles obtained with the Bauer-Kirby technique. Approximately 4,000 microorganisms, routinely identified and tested for antibiotic sensitivities in a large hospital microbiology laboratory, were used as a data set for several pattern recognition classification methods: K--nearest-neighbor analysis, statistical isolinear multicomponent analysis, Bayesian inference, and linear discriminant analysis. K--nearest-neighbor analysis yielded the highest prospective classification accuracy for gram-negative organisms, 90%. When those organisms displaying an atypical antibiotic resistance pattern were excluded from the data, the gram-negative classification accuracy improved to 95%. These results are inferior to currently accepted biochemical identification methods. Microorganisms with atypical antibiotic resistance patterns are likely to be misidentified and are common enough (17% of our isolates) to limit the feasibility of routine identification of microorganisms from their antibiotic sensitivities.

Anti-Bacterial Agents