Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Machine learning in bioinformatics.

This article reviews machine learning methods for bioinformatics. It presents modelling methods, such as supervised classification, clustering and probabilistic graphical models for knowledge discovery, as well as deterministic and stochastic heuristics for optimization. Applications in genomics, proteomics, systems biology, evolution and text mining are also shown.

Artificial Intelligence↗

Asymptotic judgment of cause in a relative validity paradigm.

We report three experiments in which we tested asymptotic and dynamic predictions of the Rescorla-Wagner (R-W) model and the asymptotic predictions of Cheng's probabilistic contrast model (PCM) concerning judgments of causality when there are two possible causal candidates. We used a paradigm in which the presence of a causal candidate that is highly correlated with an effect influences judgments of a second, moderately correlated or uncorrelated cause. In Experiment 1, which involved a moderate outcome density, judgments of a moderately positive cause were attenuated when it was paired with either a perfect positive or perfect negative cause. This attenuation was robust over a large set of trials but was greater when the strong predictor was positive. In Experiment 2, in which there was a low overall density of outcomes, judgments of a moderately correlated positive cause were elevated when this cause was paired with a perfect negative causal candidate. This elevation was also quite robust over a large set of trials. In Experiment 3, estimates of the strength of a causal candidate that was uncorrelated with the outcome were reduced when it was paired with a perfect cause. The predictions of three theoretical models of causal judgments are considered. Both the R-W model and Cheng's PCM accounted for some but not all aspects of the data. Pearce's model of stimulus generalization accounts for a greater proportion of the data.

Adult↗

A central spectrum model: a synthesis of auditory-nerve timing and place cues in monaural communication of frequency spectrum.

A probabilistic psychophysical model for monaural communication from the auditory nerve to the brain is given in the form of a tonotopic display of stimulus spectrum, termed central spectrum. The model builds upon prior research demonstrating the potential of neural timing cues from the auditory nerve for conveying information on complex spectra, and was designed to meet the quantified demands of the psychophysics of frequency measurement. The central spectrum magnitude at each frequency is determined by the response of the auditory-nerve fibre with characteristic frequency matching that frequency. An interval histogram from each fiber is passed through a filter matched to the characteristic frequency of the fiber. This output versus characteristic frequency defines the central spectrum. Detailed analysis demonstrates that efficient probabilistic processing of the central spectrum described known psychophysical properties of frequency measurement in discrimination and periodicity pitch experiments. Psychophysical models based upon the central spectrum model followed by optimum probabilistic pattern recognition are potentially relevant for predicting human communication limits in response to arbitrary sounds of speech and music.

Cues↗

A probabilistic neural network approach for modeling and classification of bacterial growth/no-growth data.

In this paper, we propose to use probabilistic neural networks (PNNs) for classification of bacterial growth/no-growth data and modeling the probability of growth. The PNN approach combines both Bayes theorem of conditional probability and Parzen's method for estimating the probability density functions of the random variables. Unlike other neural network training paradigms, PNNs are characterized by high training speed and their ability to produce confidence levels for their classification decision. As a practical application of the proposed approach, PNNs were investigated for their ability in classification of growth/no-growth state of a pathogenic Escherichia coli R31 in response to temperature and water activity. A comparison with the most frequently used traditional statistical method based on logistic regression and multilayer feedforward artificial neural network (MFANN) trained by error backpropagation was also carried out. The PNN-based models were found to outperform linear and nonlinear logistic regression and MFANN in both the classification accuracy and ease by which PNN-based models are developed.

Bacteria↗

Atypical squamous cells of undetermined significance in liquid-based cytologic specimens: results of reflex human papillomavirus testing and histologic follow-up in routine practice with comparison of interpretive and probabilistic reporting methods.

BACKGROUND: Human papillomavirus (HPV) DNA testing for high-risk types after Papanicolaou (Pap) smear interpretations of atypical squamous cells of undetermined significance (ASCUS) is a sensitive method for identifying women who harbor underlying high-grade squamous intraepithelial lesions (HSIL). To the authors' knowledge, the application of HPV testing to ASCUS smears in routine practice with comparison of probabilistic and interpretive models of cytologic reporting has not been reported. METHODS: HPV DNA testing was performed reflexively on 216 liquid-based Pap smears that initially were interpreted as ASCUS. According to the interpretive model, ASCUS interpretations were modified and reported as either low-grade squamous intraepithelial lesions (LSIL) or squamous intraepithelial lesions (SIL) when HPV positive and as reactive when HPV negative. Using the probabilistic model, ASCUS interpretations were maintained and simply reported with the HPV test result. Histologic follow-up data were obtained. RESULTS: Of the 216 women with ASCUS cytology, 142 (65.7%) were positive for high-risk HPV types. Of the 142 HPV-positive ASCUS smears, 101 (71.1%) were modified to an interpretation of LSIL (96 cases) or SIL (5 cases). Histologic follow-up of 55 of the 101 HPV-positive smears in the interpretive group and 26 of the 41 HPV-positive smears in the probabilistic group yielded similar percentages of lesions (18 lesions [32.7%] and 9 lesions [34.6%], respectively). However, there was a preponderance of low-grade lesions in the interpretive group (89%) but a nearly equal distribution of low-grade and high-grade lesions in the probabilistic group (56% and 44%, respectively); overall, 22% of the lesions were high-grade. Of the 74 HPV-negative ASCUS smears, 71 (96%) were modified to reactive and all 5 with histologic follow-up were judged as negative. CONCLUSIONS: Colposcopy with tissue studies was virtually restricted to HPV-positive cases, regardless of the reporting model used, suggesting that clinicians are basing colposcopy triage on the HPV test result rather than the definitiveness of the cytologic interpretation. This observation, the similar yield of lesions in both groups, and the significant risk of high-grade lesions argue against application of the interpretive model to HPV-tested ASCUS cases.

Adolescent↗

A marketing orientation to modeling the hospital-supplier interface: a probabilistic approach.

The author adopts a marketing orientation to model the hospital-supplier interface. A probabilistic approach using logit models is employed. Internal validity of the models estimated is examined and found to be satisfactory. The implications of the modeling process and findings for both hospitals and linen service contractors are discussed. The study reported is the second in a programmatic inquiry. The results of the first study were reported in the March 1986 issue of this journal.

Bedding and Linens↗

Probabilistic formalism and hierarchy of models for polydispersed turbulent two-phase flows.

This paper deals with a probabilistic approach to polydispersed turbulent two-phase flows following the suggestions of Pozorski and Minier [Phys. Rev. E 59, 855 (1999)]. A general probabilistic formalism is presented in the form of a two-point Lagrangian PDF (probability density function). A new feature of the present approach is that both phases, the fluid as well as the particles, are included in the PDF description. It is demonstrated how the formalism can be used to show that there exists a hierarchy between the classical approaches such as the Eulerian and Lagrangian methods. It is also shown that the Eulerian and Lagrangian models can be obtained in a systematic way from the PDF formalism. Connections with previous papers are discussed.

Journal Article↗

Probabilistic sensitivity analysis incorporating the bootstrap: an example comparing treatments for the eradication of Helicobacter pylori.

Decision-analytic models are frequently used to evaluate the relative costs and benefits of alternative therapeutic strategies for health care. Various types of sensitivity analysis are used to evaluate the uncertainty inherent in the models. Although probabilistic sensitivity analysis is more difficult theoretically and computationally, the results can be much more powerful and useful than deterministic sensitivity analysis. The authors show how a Monte Carlo simulation can be implemented using standard software to perform a probabilistic sensitivity analysis incorporating the bootstrap. The method is applied to a decision-analytic model evaluating the cost-effectiveness of Helicobacter pylori eradication. The necessary steps are straightforward and are described in detail. The use of the bootstrap avoids certain difficulties encountered with theoretical distributions. The probabilistic sensitivity analysis provided insights into the decision-analytic model beyond the traditional base-case and deterministic sensitivity analyses and should become the standard method for assessing sensitivity.

Anti-Bacterial Agents↗

Probabilities and polarity biases in conditional inference.

A probabilistic computational level model of conditional inference is proposed that can explain polarity biases in conditional inference (e.g., J. St. B. T. Evans, 1993). These biases are observed when J. St. B. T. Evans's (1972) negations paradigm is used in the conditional inference task. The model assumes that negations define higher probability categories than their affirmative counterparts (M. Oaksford & K. Stenning, 1992); for example, P(not-dog) > P(dog). This identification suggests that polarity biases are really a rational effect of high-probability categories. Three experiments revealed that, consistent with this probabilistic account, when high-probability categories are used instead of negations, a high-probability conclusion effect is observed. The relationships between the probabilistic model and other phenomena and other theories in conditional reasoning are discussed.

Humans↗

Spray irrigation of landfill leachate: estimating potential exposures to workers and bystanders using a modified air box model and generalised source term.

Generalised source term data from UK leachates and a probabilistic exposure model (BPRISC(4)) were used to evaluate key routes of exposure from chemicals of concern during the spraying irrigation of landfill leachate. Risk estimates secured using a modified air box model are reported for a hypothetical worker exposed to selected chemicals within a generalised conceptual exposure model of spray irrigation. Consistent with pesticide spray exposure studies, the key risk driver is dermal exposure to the more toxic components of leachate. Changes in spray droplet diameter (0.02-0.2 cm) and in spray flow rate (50-1000 l/min) have little influence on dermal exposure, although the lesser routes of aerosol ingestion and inhalation are markedly affected. The risk estimates modelled using this conservative worst case exposure scenario are not of sufficient magnitude to warrant major concerns about chemical risks to workers or bystanders from this practice in the general sense. However, the modelling made use of generic concentration data for only a limited number of potential landfill leachate contaminants, such that individual practices may require assessment on the basis of their own merits.

Aerosols↗

Two item response theory models for analysing normative forced-choice personality items.

This paper proposes two unidimensional item response theory (IRT) models for analysing normative forced-choice personality items. Both models are derived from a common theoretical framework and arise as a result of different assumptions regarding the mechanism of choice. The simplest mechanism gives rise to the one-parameter normal-ogive model. The second mechanism gives rise to a new IRT model, which is closely related to the Coombs-Zinnes probabilistic unfolding model. The second model is compared theoretically to the normal-ogive model in terms of item characteristic curves and amount of item information. Next, procedures for estimating the respondent and the item parameters in the second model are described. Finally, both models are empirically compared by using two well-known personality measures.

Arousal↗

Probabilistic analysis and computationally expensive models: Necessary and required?

OBJECTIVE: To assess the importance of considering decision uncertainty, the appropriateness of probabilistic sensitivity analysis (PSA), and the use of patient-level simulation (PLS) in appraisals for the National Institute for Health and Clinical Excellence (NICE). METHODS: Decision-makers require estimates of decision uncertainty alongside expected net benefits (NB) of interventions. This requirement may be difficult in computationally expensive models, for example, those employing PLS. NICE appraisals published up until January 2005 were reviewed to identify those where the assessment group utilized a PLS model structure to estimate NB. After identifying PLS models, all appraisals published in the same year were reviewed. RESULTS: Among models using PLS, one out of six conducted PSA, compared with 16 out of 24 cohort models. Justification for omitting PSA was absent in most cases. Reasons for choosing PLS included treatment switching, sampling patient characteristics and dependence on patient history. Alternative modeling approaches exist to handle these, including semi-Markov models and emulators that eliminate the need for two-level simulation. Stochastic treatment switching and sampling baseline characteristics do not inform adoption decisions. Modeling patient history does not necessitate PLS, and can depend on the software used. PLS addresses nonlinear relationships between patient variability and model outputs, but other options exist. Increased computing power, emulators or closed-form approximations can facilitate PSA in computationally expensive models. CONCLUSIONS: In developing models analysts should consider the dual requirement of estimating expected NB and characterizing decision uncertainty. It is possible to develop models that meet these requirements within the constraints set by decision-makers.

Computer Simulation↗

CONTRAfold: RNA secondary structure prediction without physics-based models.

MOTIVATION: For several decades, free energy minimization methods have been the dominant strategy for single sequence RNA secondary structure prediction. More recently, stochastic context-free grammars (SCFGs) have emerged as an alternative probabilistic methodology for modeling RNA structure. Unlike physics-based methods, which rely on thousands of experimentally-measured thermodynamic parameters, SCFGs use fully-automated statistical learning algorithms to derive model parameters. Despite this advantage, however, probabilistic methods have not replaced free energy minimization methods as the tool of choice for secondary structure prediction, as the accuracies of the best current SCFGs have yet to match those of the best physics-based models. RESULTS: In this paper, we present CONTRAfold, a novel secondary structure prediction method based on conditional log-linear models (CLLMs), a flexible class of probabilistic models which generalize upon SCFGs by using discriminative training and feature-rich scoring. In a series of cross-validation experiments, we show that grammar-based secondary structure prediction methods formulated as CLLMs consistently outperform their SCFG analogs. Furthermore, CONTRAfold, a CLLM incorporating most of the features found in typical thermodynamic models, achieves the highest single sequence prediction accuracies to date, outperforming currently available probabilistic and physics-based techniques. Our result thus closes the gap between probabilistic and thermodynamic models, demonstrating that statistical learning procedures provide an effective alternative to empirical measurement of thermodynamic parameters for RNA secondary structure prediction. AVAILABILITY: Source code for CONTRAfold is available at http://contra.stanford.edu/contrafold/.

Algorithms↗

Testing of metal bioaccumulation models with measured body burdens in mice.

Estimates of chemical accumulation in prey organisms can contribute considerable uncertainty to predictive ecological risk assessments. Comparing body burdens calculated in food web models with measured tissue concentrations provides essential information about the expected accuracy of risk indices. Estimates of arsenic, cadmium, copper, lead, and nickel body burdens in house mice (Mus musculus) inhabiting a seasonal wetland were generated with two small mammal bioaccumulation models. Published soil-to-small mammal bioaccumulation regression models produced accurate estimates of arsenic and lead body burdens but failed to adequately predict copper and nickel levels in mice. Incorporating conservative prediction intervals in the regression models shows potential for successful applications in screening-level risk assessments. A simple mechanistic cumulative ingestion bioaccumulation model overpredicted lead levels in mice generally by less than one order of magnitude but greatly overpredicted concentrations of arsenic, copper, and nickel. Better estimates of absorption and elimination of ingested metals and knowledge of specific arthropod taxa in house mouse diets are likely to improve the accuracy of the cumulative ingestion model. Applying Monte Carlo simulations to the soil-small mammal regression models generated probabilistic estimates of body burdens that were consistent with deterministic results. However, deterministic minimum and maximum predictions of the ingestion model were excessively conservative (widely spaced) relative to lower and upper probabilistic percentiles. Metal levels predicted in individual mice on the basis of mouse-specific parameter values and exposures were not significantly more accurate than bioaccumulation predictions for the sitewide population.

Animals↗

Identification Model Based on the Maximum Information Entropy Principle.

A new theoretical approach to stimulus identification is proposed through a probabilistic multidimensional model based on the maximum information entropy principle. The approach enables us to derive the multidimensional scaling (MDS) choice model, without appealing to Luce's choice rule and without defining a similarity function. It also clarifies the relationship between the MDS choice model and the optimal version of the identification model based on Ashby's general recognition theory; it is shown theoretically that the identification model derived from the new approach includes these two models as special cases. Finally, as an application of our approach, a model of similarity judgment is proposed and compared with Ashby's extended similarity model. Copyright 2001 Academic Press.

Journal Article↗

Stochastic models of soil denitrification.

Soil denitrification is a highly variable process that appears to be lognormally distributed. This variability is manifested by large sample coefficients of variation for replicate estimates of soil core denitrification rates. Deterministic models for soil denitrification have been proposed in the past, but none of these models predicts the approximate lognormality exhibited by natural denitrification rate estimates. In this study, probabilistic (stochastic) models were developed to understand how positively skewed distributions for field denitrification rate estimates result from the combined influences of variables known to affect denitrification. Three stochastic models were developed to describe the distribution of measured soil core denitrification rates. The driving variables used for all the models were denitrification enzyme activity and CO(2) production rates. The three models were distinguished by the functional relationships combining these driving variables. The functional relationships used were (i) a second-order model (model 1), (ii) a second-order model with a threshold (model 2), and (iii) a second-order saturation model (model 3). The parameters of the models were estimated by using 12 separate data sets (24 replicates per set), and their abilities to predict denitrification rate distributions were evaluated by using three additional independent data sets of 180 replicates each. Model 2 was the best because it produced distributions of denitrification rate which were not significantly different (P > 0.1) from distributions of measured denitrification rates. The generality of this model is unknown, but it accurately predicted the mean denitrification rates and accounted for the stochastic nature of this variable at the site studied. The approach used in this study may be applicable to other areas of ecological research in which accounting for the high spatial variability of microbiological processes is of interest.

Journal Article↗

Use of cell proliferation data in modeling urinary bladder carcinogenesis.

A multistage, probabilistic, biologically based model of carcinogenesis has been developed involving qualitative and quantitative aspects of the process. A chemical can affect the risk of cancer by directly damaging DNA and/or increasing the number of cell divisions during which errors in DNA replication can occur. Based on this model, carcinogens are classified as genotoxic versus nongenotoxic; nongenotoxic chemicals are further divided on the basis of whether or not they act through a specific cell receptor. Nongenotoxic compounds, particularly those acting through a nonreceptor mechanism, are likely to have dose and/or species-specific thresholds. This classification also implies the existence of chemicals that will be carcinogenic at high doses in animal models, but because of dose and/or mechanistic considerations, will not be carcinogenic to humans at levels of exposure. N-[4-(5-nitro-2-furyl)-2-thiazolyl] formamide (FANFT) and 2-acetylaminofluorene (AAF) are classical genotoxic bladder carcinogens that also cause proliferative effects at higher doses. Although there is an apparent no-effect level for the urinary bladder carcinogenicity of these two compounds at low doses, in reality, DNA adducts form at these low levels, and it is likely that there is a cancer effect (no threshold), but it is below the level of detection of the bioassay. These conclusions are based on studies involving multiple doses and time points in rodents, including results from the ED01. Pellets implanted directly into the rodent bladder lumen or calculi formed in the urine as a result of an administered chemical cause abrasion of the urothelium, and a marked increase in cell proliferation and cell number, and ultimately tumors.(ABSTRACT TRUNCATED AT 250 WORDS)

2-Acetylaminofluorene↗

A graphical model for estimating stimulus-evoked brain responses from magnetoencephalography data with large background brain activity.

This paper formulates a novel probabilistic graphical model for noisy stimulus-evoked MEG and EEG sensor data obtained in the presence of large background brain activity. The model describes the observed data in terms of unobserved evoked and background factors with additive sensor noise. We present an expectation maximization (EM) algorithm that estimates the model parameters from data. Using the model, the algorithm cleans the stimulus-evoked data by removing interference from background factors and noise artifacts and separates those data into contributions from independent factors. We demonstrate on real and simulated data that the algorithm outperforms benchmark methods for denoising and separation. We also show that the algorithm improves the performance of localization with beamforming algorithms.

Algorithms↗