Search PubMedSearch

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

TreeFlow: Probabilistic Modelling and Automatic Differentiation for Phylogenetics.

Probabilistic modelling frameworks are powerful tools for statistical modelling and inference. They are not immediately generalizable to phylogenetic problems due to the particular computational properties of the phylogenetic tree object. TreeFlow is a software library for probabilistic modelling and automatic differentiation with phylogenetic trees. It embeds phylogenetic trees in the TensorFlow Probability framework, and implements inference algorithms for phylogenetic models given a fixed tree topology. We demonstrate how TreeFlow can be used to quickly implement and assess new models. We also show that it provides reasonable performance for gradient-based inference algorithms compared to specialized computational libraries for phylogenetics.

Bayesian inference

Probabilistic diagnosis using a reformulation of the INTERNIST-1/QMR knowledge base. I. The probabilistic model and inference algorithms.

In Part I of this two-part series, we report the design of a probabilistic reformulation of the Quick Medical Reference (QMR) diagnostic decision-support tool. We describe a two-level multiply connected belief-network representation of the QMR knowledge base of internal medicine. In the belief-network representation of the QMR knowledge base, we use probabilities derived from the QMR disease profiles, from QMR imports of findings, and from National Center for Health Statistics hospital-discharge statistics. We use a stochastic simulation algorithm for inference on the belief network. This algorithm computes estimates of the posterior marginal probabilities of diseases given a set of findings. In Part II of the series, we compare the performance of QMR to that of our probabilistic system on cases abstracted from continuing medical education materials from Scientific American Medicine. In addition, we analyze empirically several components of the probabilistic model and simulation algorithm.

Algorithms

Mutation and childhood cancer: a probabilistic model for the incidence of retinoblastoma.

The incidences of some childhood cancers have been shown to fit a two-mutation hypothesis for cancer initiation. According to this hypothesis, the first mutation can be either germinal or somatic while the second is always somatic. A probabilistic model involving the mean number of tumors per genetically susceptible individual is developed as a function of age and is compared with age incidence data for retinoblastoma. The change in the mean number of tumors with time is interpreted in terms of the growth of retinal cells. In patients who are not genetically susceptible, the times of occurrence of the first and second somatic mutations can be inferred from a comparison of familial and non-familial unilateral case incidences. The total incidences of hereditary and nonhereditary forms of retinoblastoma are related to germinal and somatic mutation rates. The even distribution of certain childhood cancers throughout the world suggests that their incidences are determined by spontaneous mutation rates rather than by local environmental mutagenic carcinogens.

Child

A probabilistic model for genetic recombination of nonreplicating lambda-phage DNA, stimulated by "mismatch repair" of UV photoproducts.

Genetic recombination of nonreplicating phage lambda-DNA, during infection of homoimmune lysogenic bacteria, was previously observed to be dramatically stimulated by prior uv irradiation of the phages, even when the Escherichia coli hosts lacked the major uv-photo-product excision-repair system (UvrABC). UvrABC-independent recombination of circular phage molecules depends on host MutHLS functions and on undermethylation of adenines at GATC sites in the phage DNA, and thus appears to be the result of "mismatch repair" of uv photoproducts. Recombinant frequencies pass through a relatively sharp maximum at 20 J/m2 and decrease at higher doses, whereas most plausible models for the process predict monotonic increases with dose, or a plateau at high uv doses. A uv-dose-dependent loss of biological activity (restriction) of all intracellular phage DNA was also observed previously. In order to provide a framework for testing possible explanations for the unusual recombinant-frequency vs uv-dose curve, a statistical model was constructed. This model includes probability terms for all possible one-exchange and two-exchange recombination processes, and incorporates the assumption that dimer recombinants are more susceptible to restriction than monomer parents (or recombinants), because of their larger target size. By adjustment of model parameters, particularly epsilon, the efficiency per photoproduct of initiation of a recombinational exchange, a theoretical dose-response curve that agreed well with experiment was obtained. The best fit corresponded to epsilon = 0.035, close to the previously observed restriction efficiency of 0.053. In the calculations, the value for h0, the average length of heteroduplex DNA, was taken to be 0.5 lambda units, i.e., about 25 kilobase pairs. This estimate for h0 was obtained here by analysis of the density distributions of the progeny of crosses between nonreplicating density-labeled lambda-phage chromosomes, published by others [M. S. Fox, C. S. Dudney and E. J. Sodergren (1979) Cold Spring Harbor Symposium on Quantitative Biology, Vo. 43, pp. 999-1007].

Bacteriophage lambda

A probabilistic model of bathing beach safety.

An improved mathematical model for bathing beach safety is proposed. It is derived by joining the probability of infection from a given dose (Poisson distribution and the probability of acquiring such a dose (lognormal distribution). Even in the absence of better clinical and epidemiological data, the model permits an assessment of relative risk from certain hazards and the design of more meaningful bacteriological standards for individual beaches.

Bacteria

Micrometastases formation: a probabilistic model.

A mathematical model of the process of metastases is formulated in which the hematogenous metastatic process from a solid tumor is considered to consist of a series of stages. A mathematical expression is obtained for the probability that no metastases will have been established by a characteristic time interval after tumor initiation. The murine T241 fibrosarcoma that rapidly and reproduceably produces pulmonary metastases was studied. Estimates of parameters required for the expression of probability of metastases formation were derived experimentally. The probability remains close to one for a characteristic time at which point it drops to zero. This indicates that at least in this experimental system there is a predictable critical time period beyond which micrometastases are virtually certain to have been formed.

Animals

Probabilistic mental models: a Brunswikian theory of confidence.

Research on people's confidence in their general knowledge has to date produced two fairly stable effects, many inconsistent results, and no comprehensive theory. We propose such a comprehensive framework, the theory of probabilistic mental models (PMM theory). The theory (a) explains both the overconfidence effect (mean confidence is higher than percentage of answers correct) and the hard-easy effect (overconfidence increases with item difficulty) reported in the literature and (b) predicts conditions under which both effects appear, disappear, or invert. In addition, (c) it predicts a new phenomenon, the confidence-frequency effect, a systematic difference between a judgment of confidence in a single event (i.e., that any given answer is correct) and a judgment of the frequency of correct answers in the long run. Two experiments are reported that support PMM theory by confirming these predictions, and several apparent anomalies reported in the literature are explained and integrated into the present framework.

Adult

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis

Deficiency in POLE Exonuclease Causes Synthetic Lethality in Highly Aneuploid Cancer Cells.

UNLABELLED: Aneuploidy is a hallmark of cancer and is associated with drug resistance and poor clinical outcomes across diverse cancer types. However, no therapies have been clinically established to target highly aneuploid tumors. By analyzing nearly half a million tumor samples subjected to comprehensive genomic profiling, we identified a striking mutual exclusivity between POLE exonuclease domain mutations and high aneuploidy burden. This observation was independently validated using data from The Cancer Genome Atlas (TCGA) and the Cancer Cell Line Encyclopedia (CCLE). Probabilistic modeling revealed that the elevated quantity and unique spectrum of mutations induced by POLE exonuclease deficiency increase the likelihood of inactivating essential genes on chromosome arms harboring losses, leading to a synthetic lethal phenotype in highly aneuploid cells. Functional experiments demonstrated that POLE exonuclease activity is essential for the viability of highly aneuploid cancer cell lines but dispensable in diploid cells. These findings suggest that selective inhibition of POLE exonuclease activity may represent a promising therapeutic strategy for targeting highly aneuploid tumors. SIGNIFICANCE: An integrated approach using large-scale genomic analyses, probabilistic modeling and functional validation identified POLE exonuclease as a potential synthetic lethal target to overcome cancer aneuploidy.

Humans

A probabilistic generative model for quantification of DNA modifications enables analysis of demethylation pathways.

We present a generative model, Lux, to quantify DNA methylation modifications from any combination of bisulfite sequencing approaches, including reduced, oxidative, TET-assisted, chemical-modification assisted, and methylase-assisted bisulfite sequencing data. Lux models all cytosine modifications (C, 5mC, 5hmC, 5fC, and 5caC) simultaneously together with experimental parameters, including bisulfite conversion and oxidation efficiencies, as well as various chemical labeling and protection steps. We show that Lux improves the quantification and comparison of cytosine modification levels and that Lux can process any oxidized methylcytosine sequencing data sets to quantify all cytosine modifications. Analysis of targeted data from Tet2-knockdown embryonic stem cells and T cells during development demonstrates DNA modification quantification at unprecedented detail, quantifies active demethylation pathways and reveals 5hmC localization in putative regulatory regions.

5-Methylcytosine

Probability of conduction deficit as related to fiber length in random-distribution models of peripheral neuropathies.

This paper presents a set of probabilistic models which reproduce the proximodistal gradient of sensory deficit in peripheral neuropathies, on the basis of the occurrence of axonal dysfunction as a result of randomly distributed abnormalities. The models, which are based on conduction block, loss of temporal coherence, and weak interactions between nerve fibers, demonstrate that randomly distributed axonal dysfunction provides a sufficient condition for distal sensory deficit. The models predict a marked reduction in the length for normal sensory conduction with small increases in the probability of axomal dysfunction, providing a possible correlate for the rapid clinical progression of some neuropathies. The hypothesis that weak interactions between fibers result in paresthesiae in peripheral neuropathies is also discussed.

Humans

Prototype of simulation models for epizootics in domestic animals.

Based on the Reed-Frost model (Model I), the authors conducted computer simulation of an epizootic model (Model II) constructed on the assumption that any infected animal in a group, after a given time-period of infectivity, would be removed from the group at the beginning of the next time-period. Models I and II were simulated 100 times for each of the different conditions, viz. the initial size of group, 100 and 1,000, the five steps of contact rate or contact size, and the five more steps of contact rate for the group of 1,000 animals in Model I. From the results obtained, it is believed that as a constant parameter, contact size may be preferably used instead of contact rate in these models. Model II mostly gave higher morbidities than Model I, and earlier termination of epizootics, except the simulation with the smallest contact size. This fact may be due to the effect of herd immunity involved only in Model I. The long duration of epizootic was demonstrated in two of the 100 simulations of Model II with 1,000 individuals and contact size 1. This is characteristic of probabilistic models which are really instructive to studying the flow of epizootic.

Animals