Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

MR-determined metabolic phenotype of breast cancer in prediction of lymphatic spread, grade, and hormone status.

The purpose of the study was to evaluate the use of metabolic phenotype, described by high-resolution magic angle spinning magnetic resonance spectroscopy (HR MAS MRS), as a tool for prediction of histological grade, hormone status, and axillary lymphatic spread in breast cancer patients. Biopsies from breast cancer (n = 91) and adjacent non-involved tissue (n = 48) were excised from patients (n = 77) during surgery. HR MAS MR spectra of intact samples were acquired. Multivariate models relating spectral data to histological grade, lymphatic spread, and hormone status were designed. The multivariate methods applied were variable reduction by principal component analysis (PCA) or partial least-squares regression-uninformative variable elimination (PLS-UVE), and modelling by PLS, probabilistic neural network (PNN), or cascade correlation neural network. In the end, model verification by prediction of blind samples (n = 12) was performed. Validation of PNN training resulted in sensitivity and specificity ranging from 83 to 100% for all predictions. Verification of models by blind sample testing showed that hormone status was well predicted by both PNN and PLS (11 of 12 correct), lymphatic spread was best predicted by PLS (8 of 12), whereas PLS-UVE PNN was the best approach for predicting grade (9 of 12 correct). MR-determined metabolic phenotype may have a future role as a supplement for clinical decision-making-concerning adjuvant treatment and the adaptation to more individualised treatment protocols.

Adult↗

Specification of models in large expert systems based on causal probabilistic networks.

Problems involved in the specification of large expert systems are discussed. In the specification of causal probabilistic networks conditional probability tables for all nodes have to be provided. These conditional probability tables can often be described by models that specify the nature of interaction between nodes. Various types of models are described and a program that handles such models is presented. Large causal probabilistic networks often contain several copies of identical tables or structures. A header facility that provides common definitions of such repeated elements is proposed. This facility makes specifications much shorter and easier to construct and maintain.

Expert Systems↗

Using fragment chemistry data mining and probabilistic neural networks in screening chemicals for acute toxicity to the fathead minnow.

The paper is illustrating how the general data mining methodology may be adapted to provide solutions to the problem of high throughput virtual screening of organic chemicals for possible acute toxicity to the fathead minnow fish. The present approach involves mining fragment information from chemical structures and is using probabilistic neural networks to model the relationship between structure and toxicity. Probabilistic neural networks implement a special class of multivariate non-linear Bayesian statistical models. The mathematical principles supporting their use for value prediction purposes are clarified and their peculiarities discussed. As part of the research phase of the data mining process, a dataset consisting of 800 structures and associated fathead minnow (Pimephales promelas) 96-h LC50 acute toxicity endpoint information is used for both the purpose of identifying an advantageous combination of fragment descriptors and for training the neural networks. As a result, two powerful models are generated. Model 1 implements the basic PNN with Gaussian kernel (statistical corrections included) while Model 2 implements the PNN with Gaussian kernel and separated variables. External validation is performed using a separate dataset consisting of 86 structures and associated toxicity information. Both learning and generalization capabilities of the two models are investigated and their limitations discussed.

Animals↗

Uncertain numbers and uncertainty in the selection of input distributions--consequences for a probabilistic risk assessment of contaminated land.

Risks from exposure to contaminated land are often assessed with the aid of mathematical models. The current probabilistic approach is a considerable improvement on previous deterministic risk assessment practices, in that it attempts to characterize uncertainty and variability. However, some inputs continue to be assigned as precise numbers, while others are characterized as precise probability distributions. Such precision is hard to justify, and we show in this article how rounding errors and distribution assumptions can affect an exposure assessment. The outcome of traditional deterministic point estimates and Monte Carlo simulations were compared to probability bounds analyses. Assigning all scalars as imprecise numbers (intervals prescribed by significant digits) added uncertainty to the deterministic point estimate of about one order of magnitude. Similarly, representing probability distributions as probability boxes added several orders of magnitude to the uncertainty of the probabilistic estimate. This indicates that the size of the uncertainty in such assessments is actually much greater than currently reported. The article suggests that full disclosure of the uncertainty may facilitate decision making in opening up a negotiation window. In the risk analysis process, it is also an ethical obligation to clarify the boundary between the scientific and social domains.

Body Weight↗

ZILA-SRM: a probabilistic framework with zero-inflated latent models for robust strain reconstruction from metagenomes.

UNLABELLED: Resolving bacterial strain diversity from shotgun metagenomic data is fundamental to understanding intra-host evolution, transmission dynamics, and phenotypic heterogeneity. However, current probabilistic approaches face a severe "identifiability limit" when disentangling highly similar genomes. Under high-noise conditions, sequencing errors, coverage overdispersion, and collinearity confound standard expectation-maximization algorithms, resulting in overfitting and spurious "ghost" strains. Here, we introduce zero-inflated latent allocation for strain reconstruction from metagenomes with adaptive sparsity regularization (ZILA-SRM) to overcome this barrier through three innovations. First, we integrate a zero-inflated Poisson mixture model to decouple "structural zeros" (true strain absence) from "sampling zeros" (stochastic dropout), addressing overdispersion in standard Poisson-based tools. Second, we impose a convex adaptive sparsity regularization penalty that leverages biological sparsity priors to shrink noise artifacts dynamically. Third, we implement a graph-theoretic refinement step using maximal clique enumeration to resolve haplotype collinearity. Benchmarking against StrainFinder and MixtureS on 702 synthetic data sets shows that ZILA-SRM achieves a 20% improvement in precision in high-complexity scenarios while maintaining over 80% recall for minor variants at 0.5% abundance. Re-analysis of deep-sequencing data from 195 Mycobacterium tuberculosis clinical samples reveals cryptic low-abundance drug-resistant variants in 12% of patients, including a minor clone carrying the rpoB S450L mutation. Furthermore, application to skin microbiome data sets further reveals a strong negative correlation between dominant Staphylococcus aureus and Staphylococcus epidermidis strains, providing genomic evidence for competitive exclusion. These findings establish ZILA-SRM as a robust tool for resolving strain-level diversity in complex metagenomes. IMPORTANCE: Understanding microbial communities at the strain level is critical because closely related strains can differ dramatically in traits such as drug resistance, virulence, and ecological interactions. However, resolving individual strains from metagenomic sequencing data remains difficult, especially when strains are highly similar or present at low abundance. As a result, biologically meaningful diversity is often obscured or misinterpreted as noise. In this study, we introduce a new framework that improves the reliability of strain reconstruction from complex metagenomic data. By reducing false-positive strain detection while preserving sensitivity to rare variants, our approach enables more accurate characterization of microbial populations. This improved resolution reveals previously hidden subpopulations in clinical and microbiome datasets, providing clearer insights into microbial evolution, competition, and the emergence of clinically relevant traits such as antibiotic resistance.

Metagenomics↗

Multi-way clustering of microarray data using probabilistic sparse matrix factorization.

MOTIVATION: We address the problem of multi-way clustering of microarray data using a generative model. Our algorithm, probabilistic sparse matrix factorization (PSMF), is a probabilistic extension of a previous hard-decision algorithm for this problem. PSMF allows for varying levels of sensor noise in the data, uncertainty in the hidden prototypes used to explain the data and uncertainty as to the prototypes selected to explain each data vector. RESULTS: We present experimental results demonstrating that our method can better recover functionally-relevant clusterings in mRNA expression data than standard clustering techniques, including hierarchical agglomerative clustering, and we show that by computing probabilities instead of point estimates, our method avoids converging to poor solutions.

Algorithms↗

Probabilistic risk assessment of abalone Haliotis diversicolor supertexta exposed to waterborne zinc.

This paper describes a risk assessment approach that integrates predicted tissue concentrations of zinc (Zn) with a concentration-response relationship and leads to predictions of survival risk for pond abalone Haliotis diversicolor supertexta as well as to the uncertainties associated with these predictions. The models implemented include a probabilistic bioaccumulation model, which linking biokinetic and consumer-resource models, accounts for Zn exposure profile and a modified Hill model for reconstructing a dose-response profile for abalone exposed to waterborne Zn. The growth risk is assessed by hazard quotients characterized by measured water level and chronic no-observed effect concentration. Our risk analyses for H. diversicolor supertexta reared near Toucheng, Kouhu, and Anping, respectively, in north, central, and south Taiwan region indicate a relatively low likelihood that survival is being affected by waterborne Zn. Expected risks of mortality for abalone were estimated as 0.46 (Toucheng), 0.36 (Kouhu), and 0.29 (Anping). The predicted 90th-percentiles of hazard quotient for potential growth risk were estimated as 1.94 (Toucheng), 0.47 (Kouhu), and 0.51 (Anping). These findings indicate that waterborne Zn exposure poses no significant risk to pond abalone in Kouhu and Anping, yet a relative high growth risk in Toucheng is alarming. Because of a scarcity of toxicity and exposure data, the probabilistic risk assessment was based on very conservative assumptions.

Animals↗

Sources of uncertainty in pesticide fate modelling.

There is worldwide interest in the application of probabilistic approaches to pesticide fate models to account for uncertainty in exposure assessments. The first steps in conducting a probabilistic analysis of any system are: (i) to identify where the uncertainties come from; and (ii) to pinpoint those uncertainties that are likely to affect most of the predictions made. This article aims at addressing those two points within the context of exposure assessment for pesticides through a review of the different sources of uncertainty in pesticide fate modelling. The extensive listing of sources of uncertainty clearly demonstrates that pesticide fate modelling is laced with uncertainty. More importantly, the review suggests that the probabilistic approaches, which are typically being deployed to account for uncertainty in the pesticide fate modelling, such as Monte Carlo modelling, ignore a number of key sources of uncertainty, which are likely to have a significant effect on the prediction of environmental concentrations for pesticides (e.g. model error, modeller subjectivity). Future research should concentrate on quantifying the impact these uncertainties have on exposure assessments and on developing procedures that enable their integration within probabilistic assessments.

Animals↗

A heterogeneous tube model of intestinal drug absorption based on probabilistic concepts.

PURPOSE: To develop an approach based on computer simulations for the study of intestinal drug absorption. METHODS: The drug flow in the gastrointestinal tract was simulated with a biased random walk model in the heterogeneous tube model (Pharm. Res. 16, 87-91, 1999), while probability concepts were used to describe the dissolution and absorption processes. An amount of drug was placed into the input end of the tube and allowed to flow, dissolve and absorb along the tube. Various drugs with a diversity in dissolution and permeability characteristics were considered. The fraction of dose absorbed (Fabs) was monitored as a function of time measured in Monte Carlo steps (MCS). The absorption number An was calculated from the mean intestinal transit time and the absorption rate constant adhering to each of the drugs examined. RESULTS: A correspondence between the probability factor used to simulate drug absorption and the conventional absorption rate constant derived from the analysis of data was established. For freely soluble drugs, the estimates for Fabs derived from simulations using as an intestinal transit time 24500 MCS (equivalent to 4.5 h) were in accord with the corresponding data obtained from literature. For sparingly soluble drugs, a comparison of the normalized concentration profiles in the tube derived from the heterogeneous tube model and the classical macroscopic mass balance approach enabled the estimation of the dissolution probability factor for five drugs examined. The prediction of Fabs can be accomplished using estimates for the absorption and the dissolution probability factors. CONCLUSIONS: A fully computerized approach which describes the flow, dissolution and absorption of drug in the gastrointestinal tract in terms of probability concepts was developed. This approach can be used to predict Fabs for drugs with various solubility and permeability characteristics provided that probability factors for dissolution and absorption are available.

Computer Simulation↗

Reformulation of thermodynamic systems with aggregation and theoretical methods for the analysis of ligand binding in proteins with monomer-multimer equilibria.

The reformulation of complex thermodynamic systems is a useful tool for their analysis as demonstrated by the theoretical analysis of conformationally mediated cooperativity in a dimeric protein. Many chemical and biochemical systems exhibit monomer-multimer equilibria, behavior not addressed in the original reformulation. A method for reformulating such systems, and the mathematical methods necessary for relating alternative models, are therefore developed. The basic principles of the reformulation are illustrated on homodimeric and heterodimeric systems. The mathematical methods necessary to relate alternative models are then derived from probabilistic considerations. Higher-order models (more interacting subunits) are related to lower-order models (fewer interacting subunits) by a polynomial expansion of the sum of species in the lower-order model to give the sum of species in the higher-order model. Using these methods, the equations describing the ligand binding behavior of a homomeric monomer-dimer system are derived. These methods are also used to relate the two alternative models for cooperativity for a homotetrameric protein; one model where the dimer is the cooperative unit and the other where the tetramer is the cooperative unit.

Ligands↗

The latent process decomposition of cDNA microarray data sets.

We present a new computational technique (a software implementation, data sets, and supplementary information are available at http://www.enm.bris.ac.uk/lpd/) which enables the probabilistic analysis of cDNA microarray data and we demonstrate its effectiveness in identifying features of biomedical importance. A hierarchical Bayesian model, called Latent Process Decomposition (LPD), is introduced in which each sample in the data set is represented as a combinatorial mixture over a finite set of latent processes, which are expected to correspond to biological processes. Parameters in the model are estimated using efficient variational methods. This type of probabilistic model is most appropriate for the interpretation of measurement data generated by cDNA microarray technology. For determining informative substructure in such data sets, the proposed model has several important advantages over the standard use of dendrograms. First, the ability to objectively assess the optimal number of sample clusters. Second, the ability to represent samples and gene expression levels using a common set of latent variables (dendrograms cluster samples and gene expression values separately which amounts to two distinct reduced space representations). Third, in constrast to standard cluster models, observations are not assigned to a single cluster and, thus, for example, gene expression levels are modeled via combinations of the latent processes identified by the algorithm. We show this new method compares favorably with alternative cluster analysis methods. To illustrate its potential, we apply the proposed technique to several microarray data sets for cancer. For these data sets it successfully decomposes the data into known subtypes and indicates possible further taxonomic subdivision in addition to highlighting, in a wholly unsupervised manner, the importance of certain genes which are known to be medically significant. To illustrate its wider applicability, we also illustrate its performance on a microarray data set for yeast.

Algorithms↗

Probabilistic assessment of biodegradability based on metabolic pathways: catabol system.

A novel mechanistic modeling approach has been developed that assesses chemical biodegradability in a quantitative manner. It is an expert system predicting biotransformation pathway working together with a probabilistic model that calculates probabilities of the individual transformations. The expert system contains a library of hierarchically ordered individual transformations and matching substructure engine. The hierarchy in the expert system was set according to the descending order of the individual transformation probabilities. The integrated principal catabolic steps are derived from set of metabolic pathways predicted for each chemical from the training set and encompass more than one real biodegradation step to improve the speed of predictions. In the current work, we modeled O2 yield during OECD 302 C (MITI I) test. MITI-I database of 532 chemicals was used as a training set. To make biodegradability predictions, the model only needs structure of a chemical. The output is given as percentage of theoretical biological oxygen demand (BOD). The model allows for identifying potentially persistent catabolic intermediates and their molar amounts. The data in the training set agreed well with the calculated BODs (r2 = 0.90) in the entire range i.e. a good fit was observed for readily, intermediate and difficult to degrade chemicals. After introducing 60% ThOD as a cut off value the model predicted correctly 98% ready biodegradable structures and 96% not ready biodegradable structures. Crossvalidation by four times leaving 25% of data resulted in Q2 = 0.88 between observed and predicted values. Presented approach and obtained results were used to develop computer software for biodegradability prediction CATABOL.

Biodegradation, Environmental↗

Identification, segmentation, and image property study of acute infarcts in diffusion-weighted images by using a probabilistic neural network and adaptive Gaussian mixture model.

RATIONALE AND OBJECTIVES: Accurate identification of infarcted regions of the brain is critical in management of stroke patients. An efficient and fast method for identification and segmentation of infarcts in the diffusion-weighted images (DWI) is proposed. MATERIALS AND METHODS: Thirteen stroke patients were studied. DWI scans were acquired with a slice thickness of 5 mm. We have used a probabilistic neural network for selecting infarct slices and an adaptive (two-level) Gaussian mixture model for segmentation of the infarcts. Statistical analysis, such as identification of distribution, first-order statistics calculation, and receiver operating characteristic curve analysis, was performed. RESULTS: The average dice index is about 0.6, and average sensitivity and specificity are about 81% and 99%, respectively. The value of sensitivity and dice index are influenced by the number of false positives and false negatives. Because artifacts and infarcts have similar imaging characteristics, it is difficult to completely eliminate the artifacts. The accuracy of localization is nearly 100% as there were only two false-positive and three false-negative slices of all 381 slices. The algorithm takes about 1 minute in the Matlab computing environment to process a volume. CONCLUSION: A method to localize and segment the acute brain infarcts is proposed. The method aids the clinician in reducing the time needed to localize and segment the infarcts. The speed of localization and segmentation can be enhanced further by implementing the algorithm in VC++ and using fast algorithms for selection of Gaussian mixture model parameters.

Algorithms↗

Assessment of human health risks for arsenic bioaccumulation in tilapia (Oreochromis mossambicus) and large-scale mullet (Liza macrolepis) from blackfoot disease area in Taiwan.

This paper carries out probabilistic risk analysis methods to quantify arsenic (As) bioaccumulation in cultured fish of tilapia (Orechromis mossambicus) and large-scale mullet (Liza macrolepis) at blackfoot disease (BFD) area in Taiwan and to assess the range of exposures for the people who eat the contaminated fish. The models implemented include a probabilistic bioaccumulation model to account for As accumulation in fish and a human health exposure and risk model that accounts for hazard quotient and lifetime risk for humans consuming contaminated fish. Results demonstrate that the ninety-fifth percentile of hazard quotient for inorganic As ranged from 0.77-2.35 for Taipei city residents with fish consumption rates of 10-70 g/d, whereas it ranged 1.86-6.09 for subsistence fishers in the BFD area with 48-143 g/d, consumption rates. The highest ninety-fifth percentile of potential health risk for inorganic As ranged from 1.92 x 10(-4)-5.25 x 10(-4) for Taipei city residents eating tilapia harvested from Hsuehchia fish farms, with consumption rates of 10-70 g/d, whereas for subsistence fishers it was 7.36 x 10(-4)-1.12 x 10(-3) with 48-143 g/d consumption rates. These findings indicate that As exposure poses risks to residents and subsistence fishers, yet these results occur under highly conservative conditions. We calculate the maximum allowable inorganic As residues associated to a standard unit risk, resulting in the maximum target residues, are 0.0019-0.0175 and 0.0023-0.0053 microg/g dry weight for tilapia and large-scale mullet, respectively, with consumption rates of 70-10 g/d, or 0.0009-0.0029 and 0.0011-0.0013 microg/g dry weight for consumption rates of 169-48 g/d.

Animals↗

The cerebellum and decision making under uncertainty.

This study aimed to identify the neural basis of probabilistic reasoning, a type of inductive inference that aids decision making under conditions of uncertainty. Eight normal subjects performed two separate two-alternative-choice tasks (the balls in a bottle and personality survey tasks) while undergoing functional magnetic resonance imaging (fMRI). The experimental conditions within each task were chosen so that they differed only in their requirement to make a decision under conditions of uncertainty (probabilistic reasoning and frequency determination required) or under conditions of certainty (frequency determination required). The same visual stimuli and motor responses were used in the experimental conditions. We provide evidence that the neo-cerebellum, in conjunction with the premotor cortex, inferior parietal lobule and medial occipital cortex, mediates the probabilistic inferences that guide decision making under uncertainty. We hypothesise that the neo-cerebellum constructs internal working models of uncertain events in the external world, and that such probabilistic models subserve the predictive capacity central to induction.

Adolescent↗

Enzymatic reactions in small spatial volumes: comment on a model of Hess and Mikhailov.

Recently Hess and Mikhailov pointed out that in small subcellular compartments diffusion is so fast that mixing is instantaneous on the time scale of many enzymatic reactions. This opens the possibility for synchronizing individual reaction events. To illustrate this fact they discuss as example an irreversible enzymatic reaction with allosteric product activation. Under appropriate conditions their model shows coherent spiking in the number of product molecules, caused by the strong correlation between reaction events. In this model only substrate binding is an indeterministic process, all other subsequent transitions between different enzyme states being deterministic, contrary to real processes. The purpose of the present paper was to investigate this interesting phenomenon by means of a more realistic modification of the original model, with only probabilistic transitions. In an attempt to obtain spiking, which was not observed under these conditions, the model was extended to make a clear distinction between allosteric high and low affinity substrate binding, in contrast to the original model using a product dependent mean binding probability. However no periodic signal was detectable in the indeterministic version of the Hess Mikhailov model or the extended version, either by means of direct visualization or on autocorrelation or Fourier analysis. Reasons why spiking is not observed in indeterministic enzyme models are discussed.

Comment↗

Probability of stimulus detection in a model population of rapidly adapting fibers.

The goal of this study is to establish a link between somatosensory physiology and psychophysics at the probabilistic level. The model for a population of monkey rapidly adapting (RA) mechanoreceptive fibers by Güçlü and Bolanowski (2002) was used to study the probability of stimulus detection when a 40 Hz sinusoidal stimulation is applied with a constant contactor size (2 mm radius) on the terminal phalanx. In the model, the detection was assumed to be mediated by one or more active fibers. Two hypothetical receptive field organizations (uniformly random and gaussian) with varying average innervation densities were considered. At a given stimulus-contactor location, changing the stimulus amplitude generates sigmoid probability-of-detection curves for both receptive field organizations. The psychophysical results superimposed on these probability curves suggest that 5 to 10 active fibers may be required for detection. The effects of the contactor location on the probability of detection reflect the pattern of innervation in the model. However, the psychophysical data do not match with the predictions from the populations with uniform or gaussian distributed receptive field centers. This result may be due to some unknown mechanical factors along the terminal phalanx, or simply because a different receptive field organization is present. It has been reported that human observers can detect one single spike in an RA fiber. By considering the probability of stimulus detection across subjects and RA populations, this article proves that more than one active fiber is indeed required for detection.

Action Potentials↗