Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Italian survey on human behaviour for inhalation exposure assessment.

In order to support risk management in identifying effective mitigation measures, exposure assessment related to environmental pollution needs to integrate monitoring of pollution levels and control data with information on population behaviour and lifestyle. With this aim, a sample population survey was carried out in a Northern Italian city, collecting data on human behavioural factors influencing inhalation exposure. Questionnaires gathering data on dwelling characteristics, and weekly individual diaries on personal behaviour, such as places frequented and daily activities, were used. Data collection was carried out in two different seasons, spring-summer and fall-winter. A sample of 270 families, randomly selected from the municipal registry, was enrolled for each seasonal observation. The study allowed quantification of variability in human behaviour revealing seasonal variation and differences due to age and gender. Daily activity patterns were described and probability distributions of inhalation rates were obtained for all observed population groups. A probabilistic exposure model was developed and the resulting exposure distributions for the two seasonal periods were compared. Results confirm that exposure estimates are strongly biased if variability in human behaviour is not taken into account.

Adult↗

[The prognostic factors and evolution of the quality of life in primary biliary cirrhosis].

The prognostic factors and the evolution of the quality of life were evaluated in 38 patients with primary biliary cirrhosis (94.7% females, mean age 52.6 +/- 2.0 years) followed up for more than 36 months (mean 65.3 +/- 3.7 months). Karnofsky's index significantly declined during follow up (p less than 0.05) in a parallel fashion to modified Child's hepatic functional class (p less than 0.05) and to the days of hospital readmission (p less than 0.05). Eleven patients (28.9%) died, and the median survival was 88.7 months. The comparison of the actuarial curves showed the following to be significant poor prognostic factors at the time of diagnosis: a) clinical: more than one associated autoimmune disease, weight loss of more than 10% of the ideal weight, jaundice, upper gastrointestinal hemorrhage associated with portal hypertension, portal-systemic encephalopathy and a modified Child's hepatic functional class of 9 or more; b) biochemical: serum albumin lower than 3.5 g/dl and bilirubin higher than 2 mg/dl; c) histological: Total histological activity index of 10 or more and erosive necrosis index of 2 or more (Knodell et al.), lobular granulomas, and stage IV (Ludwig et al). A significant correlation was found (p less than 0.001) between the R index of the Mayo Clinic and the mean survival time of our patients. As a temporary policy, we indicate hepatic transplant when R is 9.2 or higher (life expectancy lower than 24 months), awaiting our own probabilistic prognostic model with the inclusion of quality of life criteria.

Adult↗

[D.lgsl. 625/1994--Protection against carcinogenic agents: the obligation to educate].

According to act 626/1994, employers have the duty to inform and train workers and their representatives. The implementation of training activities requires the following points: planning the training program according to the needs of the target population, use of the methods aimed at promoting learning and the adoption of safe behaviour, setting-up of evaluation tools. The disciplines of risk perception and communication and adult training may provide useful contribution in this frame. At the light of the preliminary experiences in this field, the importance of the following items for workers, workers representatives and employers is emphasized: probabilistic causality models, role of cognitive and emotional factors in the learning process, definition of carcinogenic according to national and international organisation, meaning of TLV with respect to carcinogenic exposure, interaction between carcinogens in the case of multiple exposition, risk evaluation, preventive measures, transfer of carcinogen risk from workplace to domestic environment, due to lack of compliance with basic hygienic rules such proper use of work clothes.

Adult↗

MutBERT: probabilistic genome representation improves genomics foundation models.

MOTIVATION: Understanding the genomic foundation of human diversity and disease requires models that effectively capture sequence variation, such as single nucleotide polymorphisms (SNPs). While recent genomic foundation models have scaled to larger datasets and multi-species inputs, they often fail to account for the sparsity and redundancy inherent in human population data, such as those in the 1000 Genomes Project. SNPs are rare in humans, and current masked language models (MLMs) trained directly on whole-genome sequences may struggle to efficiently learn these variations. Additionally, training on the entire dataset without prioritizing regions of genetic variation results in inefficiencies and negligible gains in performance. RESULTS: We present MutBERT, a probabilistic genome-based masked language model that efficiently utilizes SNP information from population-scale genomic data. By representing the entire genome as a probabilistic distribution over observed allele frequencies, MutBERT focuses on informative genomic variations while maintaining computational efficiency. We evaluated MutBERT against DNABERT-2, various versions of Nucleotide Transformer, and modified versions of MutBERT across multiple downstream prediction tasks. MutBERT consistently ranked as one of the top-performing models, demonstrating that this novel representation strategy enables better utilization of biobank-scale genomic data in building pretrained genomic foundation models. AVAILABILITY AND IMPLEMENTATION: https://github.com/ai4nucleome/mutBERT.

Humans↗

Probabilistic discovery of overlapping cellular processes and their regulation.

In this paper, we explore modeling overlapping biological processes. We discuss a probabilistic model of overlapping biological processes, gene membership in those processes, and an addition to that model that identifies regulatory mechanisms controlling process activation. A key feature of our approach is that we allow genes to participate in multiple processes, thus providing a more biologically plausible model for the process of gene regulation. We present algorithms to learn each model automatically from data, using only genomewide measurements of gene expression as input. We compare our results to those obtained by other approaches and show that significant benefits can be gained by modeling both the organization of genes into overlapping cellular processes and the regulatory programs of these processes. Moreover, our method successfully grouped genes known to function together, recovered many regulatory relationships that are known in the literature, and suggested novel hypotheses regarding the regulatory role of previously uncharacterized proteins.

Computational Biology↗

The sensitivity of probabilistic risk assessment results to alternative model structures: a case study of municipal waste incineration.

In this analysis, human health risk due to exposure to municipal waste incinerator emissions is assessed as an example of the application of probabilistic techniques (e.g., Monte Carlo or Latin Hypercube simulations). Incinerator risk assessments are characterized by the dominance of indirect exposure, thus this analysis focuses on exposure via the ingestion of locally grown foods. In addition, since exposure to 2,3,7,8-TCDD drives most incinerator risk assessments, this compound is the subject of the illustrative calculations. An important part of probabilistic risk assessment is determining the relative influence of the input parameters on the magnitude of the variance in the output distribution. This constitutes an important step toward prioritizing data needs for additional research. However, under various possible model forms reflecting incompletely understood aspects of contaminant transport, differences may be observed in estimates of risk, variance in risk, and the relative contributions of individual uncertain and variable inputs to the variance. In this analysis, a sequential structural decomposition of the relationships between the input variables is used to partition the variance in the output (i.e., risk) to identify the most influential contributors to overall variance among them. For comparison, the partitioning of variance is repeated, using techniques of multivariate regression. In summary, this study considers the degree to which results of a probabilistic assessment are contingent on critical model assumptions about the representation of deposition velocity. Specifically, this analysis assesses the impact on the results of uncertainty about the best model of the vapor/particle partitioning behavior of semi-volatile airborne pollutants.

Air Pollution↗

Current knowledge and recent developments in consumer exposure assessment of pesticides: a UK perspective.

Techniques employed in the assessment of consumer exposure to pesticides are currently being reviewed in the UK. This is not a formal process as is happening in the USA. However, the advent of probabilistic approaches and sophisticated computer models has prompted regulators, industry and other stakeholders in the UK to recognize the need for refinements in the risk-assessment process. Sources of information and data necessary to explore such refinements are disparate. This review aims to collate the information to present a coherent picture of the current knowledge, the data available and the stakeholders involved. It can then be used as a resource with which to investigate further more specific issues. Although focussing on the UK, the European context is included and reference is made to US models and developments that should be investigated. Factors hampering progress include the lack of sufficient data on which to base quantitative analysis, especially in the residential pesticides sector, and lack of experience in using and interpreting probabilistic models. At present, such techniques are being approached with some caution in the UK and in Europe, although their utility for cumulative assessment is accepted. Communicating results to both risk managers and consumers will be a considerable challenge.

Community Participation↗

Model-independent mean-field theory as a local method for approximate propagation of information.

We present a systematic approach to mean-field theory (MFT) in a general probabilistic setting without assuming a particular model. The mean-field equations derived here may serve as a local, and thus very simple, method for approximate inference in probabilistic models such as Boltzmann machines or Bayesian networks. Our approach is 'model-independent' in the sense that we do not assume a particular type of dependences; in a Bayesian network, for example, we allow arbitrary tables to specify conditional dependences. In general, there are multiple solutions to the mean-field equations. We show that improved estimates can be obtained by forming a weighted mixture of the multiple mean-field solutions. Simple approximate expressions for the mixture weights are given. The general formalism derived so far is evaluated for the special case of Bayesian networks. The benefits of taking into account multiple solutions are demonstrated by using MFT for inference in a small and in a very large Bayesian network. The results are compared with the exact results.

Child↗

Pollen limitation or mate search need not induce an Allee effect.

When a process modelling the availability of gametes is included explicitly in population models a critical depensation or Allee effect usually results. Non-spatial models cannot describe clumping and so small populations must be assumed very diffuse. Consequently individuals in small populations experience low contact rates and so reproduction is limited. In Nature invasions into new territory are unlikely to be as diffuse as those described by non-spatial models. We develop pair approximations to a probabilistic cellular automata model with independent pollination and seed setting processes (equivalently mate search and reproduction processes). Each process can be either global (population-wide) or local (within a small neighbourhood) or a mixture of the two. When either process is global the resulting model recaptures the Allee effect found in non-spatial models. However, if both processes are at least partially local we obtain a model in which Allee effects can be avoided altogether if individuals are suitably strong pollinators and colonisers. The Allee effect disappears because small populations are dramatically more clumped when colonisation is local and less wasteful of pollen when pollination is local.

Algorithms↗

Catastrophe loss modelling of storm-surge flood risk in eastern England.

Probabilistic catastrophe loss modelling techniques, comprising a large stochastic set of potential storm-surge flood events, each assigned an annual rate of occurrence, have been employed for quantifying risk in the coastal flood plain of eastern England. Based on the tracks of the causative extratropical cyclones, historical storm-surge events are categorized into three classes, with distinct windfields and surge geographies. Extreme combinations of "tide with surge" are then generated for an extreme value distribution developed for each class. Fragility curves are used to determine the probability and magnitude of breaching relative to water levels and wave action for each section of sea defence. Based on the time-history of water levels in the surge, and the simulated configuration of breaching, flow is time-stepped through the defences and propagated into the flood plain using a 50 m horizontal-resolution digital elevation model. Based on the values and locations of the building stock in the flood plain, losses are calculated using vulnerability functions linking flood depth and flood velocity to measures of property loss. The outputs from this model for a UK insurance industry portfolio include "loss exceedence probabilities" as well as "average annualized losses", which can be employed for calculating coastal flood risk premiums in each postcode.

Computer Simulation↗

Probabilistic independence networks for hidden Markov probability models.

Graphical techniques for modeling the dependencies of random variables have been explored in a variety of different areas, including statistics, statistical physics, artificial intelligence, speech recognition, image processing, and genetics. Formalisms for manipulating these models have been developed relatively independently in these research communities. In this paper we explore hidden Markov models (HMMs) and related structures within the general framework of probabilistic independence networks (PINs). The paper presents a self-contained review of the basic principles of PINs. It is shown that the well-known forward-backward (F-B) and Viterbi algorithms for HMMs are special cases of more general inference algorithms for arbitrary PINs. Furthermore, the existence of inference and estimation algorithms for more general graphical models provides a set of analysis tools for HMM practitioners who wish to explore a richer class of HMM structures. Examples of relatively complex models to handle sensor fusion and coarticulation in speech recognition are introduced and treated within the graphical model framework to illustrate the advantages of the general approach.

Algorithms↗

Dynamic behavior of driven interfaces in models with two absorbing states.

We study the dynamics of an interface (active domain) between different absorbing regions in models with two absorbing states in one dimension: probabilistic cellular automata models and interacting monomer-dimer models. These models exhibit a continuous transition from an active phase into an absorbing phase, which belongs to the directed Ising (DI) universality class. In the active phase, the interface spreads ballistically into the absorbing regions and the interface width diverges linearly in time. Approaching the critical point, the spreading velocity of the interface vanishes algebraically with a DI critical exponent. Introducing a symmetry-breaking field h that prefers one absorbing state over the other drives the interface to move asymmetrically toward the unpreferred absorbing region. In Monte Carlo simulations, we find that the spreading velocity of this driven interface shows a discontinuous jump at criticality. We explain that this unusual behavior is due to a finite relaxation time in the absorbing phase. The crossover behavior from the symmetric case (DI class) to the asymmetric case (directed percolation class) is also studied. We find the scaling dimension of the symmetry-breaking field y(h)=1.21(5).

Journal Article↗

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar↗

Incorporating direct and indirect evidence using bayesian methods: an applied case study in ovarian cancer.

OBJECTIVE: To demonstrate the application of a Bayesian mixed treatment comparison (MTC) model to synthesize data from clinical trials to inform decisions based on all relevant evidence. METHODS: The value of an MTC model is demonstrated using a probabilistic decision-analytic model developed to assess the cost-effectiveness of second-line chemotherapy in ovarian cancer. Three clinical trials were found that each made a different pair-wise comparison of three treatments of interest in the overall patient population. As no common comparator existed between the three trials, an MTC model was used to assess the combined weight of evidence on survival from all three trials simultaneously. This analysis was compared to an alternative approach that combined two of the trials to make the same comparison of all three treatments using a common comparator, and an informal approach that did not synthesize the available evidence. RESULTS: By including all three trials using an MTC model, the credible intervals around estimated overall survival were reduced compared with making the same comparison using only two trials and a common comparator. Nevertheless, the survival estimates from the MTC model result in greater uncertainty around the optimal treatment strategy at a cost-effectiveness threshold of 30,000 pounds per quality-adjusted life-year. CONCLUSIONS: MTC models can be used to combine more data than would typically be included in a traditional meta-analysis that relies on a common comparator. They can formally quantify the combined uncertainty from all available evidence, and can be conducted using the same analytical approaches as standard meta-analyses.

Antineoplastic Agents↗

An exact analysis of the multistage model explaining dose-response concavity.

The traditional multistage (MS) model of carcinogenesis implies several empirically testable properties for dose-response functions. These include convex (linear or upward-curving) cumulative hazards as a function of dose; symmetric effects on lifetime tumor probability of transition rates at different stages; cumulative hazard functions that increase without bound as stage-specific transition rates increase without bound; and identical tumor probabilities for individuals with identical parameters and exposures. However, for at least some chemicals, cumulative hazards are not convex functions of dose. This paper shows that none of these predicted properties is implied by the mechanistic assumptions of the MS model itself. Instead, they arise from the simplifying "rare-tumor" approximations made in the usual mathematical analysis of the model. An alternative exact probabilistic analysis of the MS model with only two stages is presented, both for the usual case where a carcinogen acts on both stages simultaneously, and also for idealized initiation-promotion experiments in which one stage at a time is affected. The exact two-stage model successfully fits bioassay data for chemicals (e.g., 1,3-butadiene) with concave cumulative hazard functions that are not well-described by the traditional MS model. Qualitative properties of the exact two-stage model are described and illustrated by least-squares fits to several real datasets. The major contribution is to show that properties of the traditional MS model family that appear to be inconsistent with empirical data for some chemicals can be explained easily if an exact, rather than an approximate model, is used. This suggests that it may be worth using the exact model in cases where tumor rates are not negligible (e.g., in which they exceed 10%). This includes the majority of bioassay experiments currently being performed.

Algorithms↗

Bringing metabolic networks to life: integration of kinetic, metabolic, and proteomic data.

BACKGROUND: Translating a known metabolic network into a dynamic model requires reasonable guesses of all enzyme parameters. In Bayesian parameter estimation, model parameters are described by a posterior probability distribution, which scores the potential parameter sets, showing how well each of them agrees with the data and with the prior assumptions made. RESULTS: We compute posterior distributions of kinetic parameters within a Bayesian framework, based on integration of kinetic, thermodynamic, metabolic, and proteomic data. The structure of the metabolic system (i.e., stoichiometries and enzyme regulation) needs to be known, and the reactions are modelled by convenience kinetics with thermodynamically independent parameters. The parameter posterior is computed in two separate steps: a first posterior summarises the available data on enzyme kinetic parameters; an improved second posterior is obtained by integrating metabolic fluxes, concentrations, and enzyme concentrations for one or more steady states. The data can be heterogeneous, incomplete, and uncertain, and the posterior is approximated by a multivariate log-normal distribution. We apply the method to a model of the threonine synthesis pathway: the integration of metabolic data has little effect on the marginal posterior distributions of individual model parameters. Nevertheless, it leads to strong correlations between the parameters in the joint posterior distribution, which greatly improve the model predictions by the following Monte-Carlo simulations. CONCLUSION: We present a standardised method to translate metabolic networks into dynamic models. To determine the model parameters, evidence from various experimental data is combined and weighted using Bayesian parameter estimation. The resulting posterior parameter distribution describes a statistical ensemble of parameter sets; the parameter variances and correlations can account for missing knowledge, measurement uncertainties, or biological variability. The posterior distribution can be used to sample model instances and to obtain probabilistic statements about the model's dynamic behaviour.

Bayes Theorem↗

Patterns of drug use among white institutionalized delinquents in Georgia: evidence from a latent class analysis.

Previous research by Kandel [1] and others indicates that adolescent drug use follows a progression from legal drugs, through marijuana, to hard drugs. In this study of drug use patterns, institutionalized delinquents were found to follow a similar progression of drug use. Unlike previous studies which used "rule of thumb" methods of model assessment, this study uses probabilistic assessment of the models. Importantly, the present study supports a modified gate-way sequence, where cocaine use appears as an intermediate step between marijuana use and use of other hard drugs. It is suggested that widespread availability of cocaine in the late 1980's may have resulted in a new "step" in the drug use sequence.

Adolescent↗

Analyzing bioterror response logistics: the case of smallpox.

To evaluate existing and alternative proposals for emergency response to a deliberate smallpox attack, we embed the key operational features of such interventions into a smallpox disease transmission model. We use probabilistic reasoning within an otherwise deterministic epidemic framework to model the 'race to trace', i.e., attempting to trace (via the infector) and vaccinate an infected person while (s)he is still vaccine-sensitive. Our model explicitly incorporates a tracing/vaccination queue, and hence can be used as a capacity planning tool. An approximate analysis of this large (16 ODE) system yields closed-form estimates for the total number of deaths and the maximum queue length. The former estimate delineates the efficacy (i.e., accuracy) and efficiency (i.e., speed) of contact tracing, while the latter estimate reveals how congestion makes the race to trace more difficult to win, thereby causing more deaths. A probabilistic analysis is also used to find an approximate closed-form expression for the total number of deaths under mass vaccination, in terms of both the basic reproductive ratio and the vaccination capacity. We also derive approximate thresholds for initially controlling the epidemic for more general interventions that include imperfect vaccination and quarantine.

Bioterrorism↗