Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Attractor dynamics in feedforward neural networks.

We study the probabilistic generative models parameterized by feedforward neural networks. An attractor dynamics for probabilistic inference in these models is derived from a mean field approximation for large, layered sigmoidal networks. Fixed points of the dynamics correspond to solutions of the mean field equations, which relate the statistics of each unit to those of its Markov blanket. We establish global convergence of the dynamics by providing a Lyapunov function and show that the dynamics generate the signals required for unsupervised learning. Our results for feedforward networks provide a counterpart to those of Cohen-Grossberg and Hopfield for symmetric networks.

Bayes Theorem↗

Gene perturbation and intervention in probabilistic Boolean networks.

MOTIVATION: A major objective of gene regulatory network modeling, in addition to gaining a deeper understanding of genetic regulation and control, is the development of computational tools for the identification and discovery of potential targets for therapeutic intervention in diseases such as cancer. We consider the general question of the potential effect of individual genes on the global dynamical network behavior, both from the view of random gene perturbation as well as intervention in order to elicit desired network behavior. RESULTS: Using a recently introduced class of models, called Probabilistic Boolean Networks (PBNs), this paper develops a model for random gene perturbations and derives an explicit formula for the transition probabilities in the new PBN model. This result provides a building block for performing simulations and deriving other results concerning network dynamics. An example is provided to show how the gene perturbation model can be used to compute long-term influences of genes on other genes. Following this, the problem of intervention is addressed via the development of several computational tools based on first-passage times in Markov chains. The consequence is a methodology for finding the best gene with which to intervene in order to most likely achieve desirable network behavior. The ideas are illustrated with several examples in which the goal is to induce the network to transition into a desired state, or set of states. The corresponding issue of avoiding undesirable states is also addressed. Finally, the paper turns to the important problem of assessing the effect of gene perturbations on long-run network behavior. A bound on the steady-state probabilities is derived in terms of the perturbation probability. The result demonstrates that states of the network that are more 'easily reachable' from other states are more stable in the presence of gene perturbations. Consequently, these are hypothesized to correspond to cellular functional states. AVAILABILITY: A library of functions written in MATLAB for simulating PBNs, constructing state-transition matrices, computing steady-state distributions, computing influences, modeling random gene perturbations, and finding optimal intervention targets, as described in this paper, is available on request from is@ieee.org.

Chromosome Mapping↗

A nonlinear model for assessing multiple probabilistic risks: a case study in South five-island of Changdao National Nature Reserve in China.

Several methods for estimating the potential impacts caused by multiple probabilistic risks have been suggested. These existing methods mostly rely on the weight sum algorithm to address the need for integrated risk assessment. This paper develops a nonlinear model to perform such an assessment. The joint probability algorithm has been applied to the model development. An application of the developed model in South five-island of Changdao National Nature Reserve, China, combining remote sensing data and a GIS technique, provides a reasonable risk assessment. Based on the case study, we discuss the feasibility of the model. We propose that the model has the potential for use in identifying the regional primary stressor, investigating the most vulnerable habitat, and assessing the integrated impact of multiple stressors.

China↗

PASS2: an automated database of protein alignments organised as structural superfamilies.

BACKGROUND: The functional selection and three-dimensional structural constraints of proteins in nature often relates to the retention of significant sequence similarity between proteins of similar fold and function despite poor sequence identity. Organization of structure-based sequence alignments for distantly related proteins, provides a map of the conserved and critical regions of the protein universe that is useful for the analysis of folding principles, for the evolutionary unification of protein families and for maximizing the information return from experimental structure determination. The Protein Alignment organised as Structural Superfamily (PASS2) database represents continuously updated, structural alignments for evolutionary related, sequentially distant proteins. DESCRIPTION: An automated and updated version of PASS2 is, in direct correspondence with SCOP 1.63, consisting of sequences having identity below 40% among themselves. Protein domains have been grouped into 628 multi-member superfamilies and 566 single member superfamilies. Structure-based sequence alignments for the superfamilies have been obtained using COMPARER, while initial equivalencies have been derived from a preliminary superposition using LSQMAN or STAMP 4.0. The final sequence alignments have been annotated for structural features using JOY4.0. The database is supplemented with sequence relatives belonging to different genomes, conserved spatially interacting and structural motifs, probabilistic hidden markov models of superfamilies based on the alignments and useful links to other databases. Probabilistic models and sensitive position specific profiles obtained from reliable superfamily alignments aid annotation of remote homologues and are useful tools in structural and functional genomics. PASS2 presents the phylogeny of its members both based on sequence and structural dissimilarities. Clustering of members allows us to understand diversification of the family members. The search engine has been improved for simpler browsing of the database. CONCLUSIONS: The database resolves alignments among the structural domains consisting of evolutionarily diverged set of sequences. Availability of reliable sequence alignments of distantly related proteins despite poor sequence identity and single-member superfamilies permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. PASS2 is accessible at http://www.ncbs.res.in/~faculty/mini/campass/pass2.html

Amino Acid Sequence↗

Support of diagnosis of liver disorders based on a causal Bayesian network model.

We describe our work on HEPAR II, a probabilistic causal model for diagnosis of liver disorders. The model, a Bayesian network capturing the causal interactions among various risk factors, diseases, symptoms, and test results, is based on expert knowledge combined with clinical data captured in medical records. The main applications of HEPAR II are assistance is diagnosis and training of beginning diagnosticians. We outline the principles of the applied approach, present a brief description of the model, and report its diagnostic performance.

Algorithms↗

Continuous-valued probabilistic behavior in a VLSI generative model.

This paper presents the VLSI implementation of the continuous restricted Boltzmann machine (CRBM), a probabilistic generative model that is able to model continuous-valued data with a simple and hardware-amenable training algorithm. The full CRBM system consists of stochastic neurons whose continuous-valued probabilistic behavior is mediated by injected noise. Integrating on-chip training circuits, the full CRBM system provides a platform for exploring computation with continuous-valued probabilistic behavior in VLSI. The VLSI CRBM's ability both to model and to regenerate continuous-valued data distributions is examined and limitations on its performance are highlighted and discussed.

Algorithms↗

Statistical significance of probabilistic sequence alignment and related local hidden Markov models.

The score statistics of probabilistic gapped local alignment of random sequences is investigated both analytically and numerically. The full probabilistic algorithm (e.g., the "local" version of maximum-likelihood or hidden Markov model method) is found to have anomalous statistics. A modified "semi-probabilistic" alignment consisting of a hybrid of Smith-Waterman and probabilistic alignment is then proposed and studied in detail. It is predicted that the score statistics of the hybrid algorithm is of the Gumbel universal form, with the key Gumbel parameter lambda taking on a fixed asymptotic value for a wide variety of scoring systems and parameters. A simple recipe for the computation of the "relative entropy," and from it the finite size correction to lambda, is also given. These predictions compare well with direct numerical simulations for sequences of lengths between 100 and 1,000 examined using various PAM substitution scores and affine gap functions. The sensitivity of the hybrid method in the detection of sequence homology is also studied using correlated sequences generated from toy mutation models. It is found to be comparable to that of the Smith-Waterman alignment and significantly better than the Viterbi version of the probabilistic alignment.

Algorithms↗

Health risk assessment on human exposed to environmental polycyclic aromatic hydrocarbons pollution sources.

To assess how the human exposure to environmental carcinogenic polycyclic aromatic hydrocarbons (PAHs) pollution sources generated from industrial, traffic and rural settings, we present a probabilistic risk model, appraised with reported empirical data. A probabilistic risk assessment framework is integrated with the potency equivalence factors (PEFs), age group-specific occupancy probability and the incremental lifetime cancer risk (ILCR) approaches to quantitatively estimate the exposure risk for three age groups of adults, children, and infants. The benzo[a]pyrene equivalents based PAH concentrations in rural, traffic, and industrial areas associated with age group-specific occupancy probability at different environmental settings are used to calculate daily exposure level through inhalation and dermal contact pathways. Risk analysis indicates that the inhalation-ILCR and dermal contact-ILCR values for adults follow a lognormal distribution with geometric mean 1.04x10(-4) and 3.85x10(-5) and geometric standard deviation 2.10 and 2.75, respectively, indicating high potential cancer risk; whereas for the infants the risk values are less than 10(-6), indicating no significant cancer risk. Sensitivity analysis indicates that the input variables of cancer slope factor and daily inhalation exposure level have the greater impact than that of body weight on the inhalation-ILCR; whereas for the dermal-ILCR, particle-bound PAH-to-skin adherence factor and daily dermal exposure level have the significant influence than that of body weight.

Adolescent↗

Incorporating quality of evidence into decision analytic modeling.

Our objective was to illustrate the effects of using stricter standards for the quality of evidence used in decision analytic modeling. We created a simple 10-parameter probabilistic Markov model to estimate the cost-effectiveness of directly observed therapy (DOT) for individuals with newly diagnosed HIV infection. We evaluated quality of evidence on the basis of U.S. Preventive Services Task Force methods, which specified 3 separate domains: study design, internal validity, and external validity. We varied the evidence criteria for each of these domains individually and collectively. We used published research as a source of data only if the quality of the research met specified criteria; otherwise, we specified the parameter by randomly choosing a number from a range within which every number has the same probability of being selected (a uniform distribution). When we did not eliminate poor-quality evidence, DOT improved health 99% of the time and cost less than 100,000 dollars per additional quality-adjusted life-year (QALY) 85% of the time. The confidence ellipse was extremely narrow, suggesting high precision. When we used the most rigorous standards of evidence, we could use fewer than one fifth of the data sources, and DOT improved health only 49% of the time and cost less than 100,000 dollars per additional QALY only 4% of the time. The confidence ellipse became much larger, showing that the results were less precise. We conclude that the results of decision modeling may vary dramatically depending on the stringency of the criteria for selecting evidence to use in the model.

CD4 Lymphocyte Count↗

Bayesian analysis, pattern analysis, and data mining in health care.

PURPOSE OF REVIEW: To discuss the current role of data mining and Bayesian methods in biomedicine and heath care, in particular critical care. RECENT FINDINGS: Bayesian networks and other probabilistic graphical models are beginning to emerge as methods for discovering patterns in biomedical data and also as a basis for the representation of the uncertainties underlying clinical decision-making. At the same time, techniques from machine learning are being used to solve biomedical and health-care problems. SUMMARY: With the increasing availability of biomedical and health-care data with a wide range of characteristics there is an increasing need to use methods which allow modeling the uncertainties that come with the problem, are capable of dealing with missing data, allow integrating data from various sources, explicitly indicate statistical dependence and independence, and allow integrating biomedical and clinical background knowledge. These requirements have given rise to an influx of new methods into the field of data analysis in health care, in particular from the fields of machine learning and probabilistic graphical models.

Bayes Theorem↗

Forward blocking depends on retrospective inferences about the presence of the blocked cue during the elemental phase.

When a compound cue AT is followed by an outcome (AT+), human participants will judge the relation between cue T and the outcome to be less strong if A alone was previously paired with the outcome (A+). According to the probabilistic contrast model, such a blocking effect is due to the fact that participants regard the A+ trials as trials on which A and the outcome are present but T is absent. The results of two studies showed that when the status of T was ambiguous during the A+ trials, judgments about T depended on subsequent information about the presence of T during the A+ trials. These findings support the probabilistic contrast model but are incompatible with the (revised) Rescorla-Wagner (Rescorla & Wagner, 1972) model.

Adult↗

A study of statistical methods for function prediction of protein motifs.

Automatic discovery of new protein motifs (i.e. amino acid patterns) is one of the major challenges in bioinformatics. Several algorithms have been proposed that can extract statistically significant motif patterns from any set of protein sequences. With these methods, one can generate a large set of candidate motifs that may be biologically meaningful. This article examines methods to predict the functions of these candidate motifs. We use several statistical methods: a popularity method, a mutual information method and probabilistic translation models. These methods capture, from different perspectives, the correlations between the matched motifs of a protein and its assigned Gene Ontology terms that characterise the function of the protein. We evaluate these different methods using the known motifs in the InterPro database. Each method is used to rank candidate terms for each motif. We then use the expected mean reciprocal rank to evaluate the performance. The results show that, in general, all these methods perform well, suggesting that they can all be useful for predicting the function of an unknown motif. Among the methods tested, a probabilistic translation model with a popularity prior performs the best.

Algorithms↗

Model-based biosignal interpretation.

Two relatively new approaches to model-based biosignal interpretation, qualitative simulation and modelling by causal probabilistic networks, are compared to modelling by differential equations. A major problem in applying a model to an individual patient is the estimation of the parameters. The available observations are unlikely to allow a proper estimation of the parameters, and even if they do, the task appears to have exponential computational complexity if the model is non-linear. Causal probabilistic networks have both differential equation models and qualitative simulation as special cases, and they can provide both Bayesian and maximum-likelihood parameter estimates, in most cases in much less than exponential time. In addition, they can calculate the probabilities required for a decision-theoretical approach to medical decision support. The practical applicability of causal probabilistic networks to real medical problems is illustrated by a model of glucose metabolism which is used to adjust insulin therapy in type I diabetic patients.

Bayes Theorem↗

The design and construction of a medical simulation model.

This paper describes the design, construction and validation of a probabilistic simulation model of patients who present with abdominal pain. The model incorporates text-book medical knowledge, clinical judgment, and statistics collected from real cases. The knowledge representation combines techniques of Bayesian network modelling with ideas of logistic discrimination. The model is shown to generate convincing, realistic cases; large numbers of artificial cases with no missing observations can be generated quickly. This should make the model a useful tool for investigating factors which limit achievable computer accuracy in the diagnosis of abdominal pain.

Abdominal Pain↗

Quantitative microbial risk assessment exemplified by Staphylococcus aureus in unripened cheese made from raw milk.

This paper discusses some of the developments and problems in the field of quantitative microbial risk assessment, especially exposure assessment and probabilistic risk assessment models. To illustrate some of the topics, an initial risk assessment was presented, in which predictive microbiology and survey data were combined with probabilistic modelling to simulate the level of Staphylococcus aureus in unripened cheese made from raw milk at the time of consumption. Due to limited data and absence of dose-response models, a complete risk assessment was not possible. Instead, the final level of bacteria was used as a proxy for the potential enterotoxin level, and thus the potential for causing illness. The assessment endpoint selected for evaluation was the probability that a cheese contained at least 6 log cfu S. aureus g(-1) at the time of consumption; the probability of an unsatisfactory cheese, P(uc). The initial level of S. aureus, followed by storage temperature had the largest influence on P(uc) at the two pH-values investigated. P(uc) decreased with decreasing pH and was up to a factor of 30 lower in low pH cheeses due to a slower growth rate. Of the model assumptions examined, i.e. the proportion of enterotoxigenic strains, the level of S. aureus in non-detect cheeses, the temperature limit for toxin production, and the magnitude and variability of the threshold for an unsatisfactory cheese, it was the latter that had the greatest impact on P(uc). The uncertainty introduced by this assumption was in most cases less than a factor of 36, the same order of magnitude as the maximum variability due to pH. Several data gaps were identified and suggestions were made to improve the initial risk assessment, which is valid only to the extent that the limited data reflected the true conditions and that the assumptions made were valid. Despite the limitations, a quantitative approach was useful to gain insights and to evaluate several factors that influence the potential risk and to make some inferences with relevance to risk management. For instance, the possible effect of using starter cultures in the cheese making process to improve the safety of these products.

Animals↗

Issues of the Human Reliability Analysis in the Context of Probabilistic Safety Studies.

This article addresses methodological issues of the human reliability analysis (HRA) in the context of probabilistic safety studies. Several conventional HRA techniques, more often used for the evaluation of the human error probabilities (HEPs), have been classified. A taxonomy of human actions, failure events, and related factors is outlined in order to distinguish action phases, human behavior types and incorrect outputs (errors of omission or commission), error types (slips, lapses, and mistakes), and performance-shaping factors (PSFs) influencing the human performance. A tree is proposed to facilitate the selection of a specific method for the evaluation of human reliability with regard to attributes of the situation analyzed. A software system based on the expert system technology to facilitate and document PSA and HRA is outlined. At the end of the article some research challenges in the domain are discussed.

expert systems↗

Probabilistic two-stage model of cell inactivation by ionizing particles.

A model of biological effects of ionizing particles, especially of protons and other ions, is proposed. The model is based on distinguishing the single-particle and collective effects of the underlying radiobiological mechanism. The probabilities of individual particles causing severe damage to DNA, their synergetic or saturation combinations, and the effect of the cellular repair system are taken into account. The model enables one to describe linear, parabolic and more complex curves, including those exhibiting low-dose hypersensitivity phenomena, in a systematic way. Global shape as well as detailed structure of survival curves might be represented, which is crucial if different fractionation schemes in radiotherapy should be assessed precisely. Experimental cell-survival data for inactivation of V79 cells by low-energy protons have been analysed and corresponding detailed characteristics of the inactivation mechanism have been derived for this case.

Animals↗