Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Modeling gene expression measurement error: a quasi-likelihood approach.

BACKGROUND: Using suitable error models for gene expression measurements is essential in the statistical analysis of microarray data. However, the true probabilistic model underlying gene expression intensity readings is generally not known. Instead, in currently used approaches some simple parametric model is assumed (usually a transformed normal distribution) or the empirical distribution is estimated. However, both these strategies may not be optimal for gene expression data, as the non-parametric approach ignores known structural information whereas the fully parametric models run the risk of misspecification. A further related problem is the choice of a suitable scale for the model (e.g. observed vs. log-scale). RESULTS: Here a simple semi-parametric model for gene expression measurement error is presented. In this approach inference is based an approximate likelihood function (the extended quasi-likelihood). Only partial knowledge about the unknown true distribution is required to construct this function. In case of gene expression this information is available in the form of the postulated (e.g. quadratic) variance structure of the data. As the quasi-likelihood behaves (almost) like a proper likelihood, it allows for the estimation of calibration and variance parameters, and it is also straightforward to obtain corresponding approximate confidence intervals. Unlike most other frameworks, it also allows analysis on any preferred scale, i.e. both on the original linear scale as well as on a transformed scale. It can also be employed in regression approaches to model systematic (e.g. array or dye) effects. CONCLUSIONS: The quasi-likelihood framework provides a simple and versatile approach to analyze gene expression data that does not make any strong distributional assumptions about the underlying error model. For several simulated as well as real data sets it provides a better fit to the data than competing models. In an example it also improved the power of tests to identify differential expression.

Calibration↗

[Study of the uni-dimensionality of the Yesavage-Brinck geriatric depression scale. Comparison between classical methods and Rasch's model].

The authors present a contribution to the french validation of the self-rating questionnaire of the depression in the elderly proposed by Yesavage and Brink (1982), the Geriatric Depression Scale (30 items). This study focusses on the assessment of the homogeneity and of the unidimensionality of this scale. 99 aged women living in old-people homes or attending a geriatric somatic day-hospital, not known to be psychiatrically ill, filled the GDS and were interviewed by either a psychiatrist or by a clinical psychologist. This interview yielded 44 cases of Major Depressive Disorder or of Dysthymia (DSM III). Firstly, we have applied the classical correlational methods of assessment of scale Reliability and Construct Validity: Cronbach's coefficient alpha and item-total correlations (homogeneity) and Principal Component Analysis (PCA) without rotation. Then, we have performed a Rasch Model Analysis: this method which belongs to the general frame of Latent Trait Theory relies on a probabilistic model of subject's response to individual questions. In the Rasch model, the response probability of a given subject to a given item is a logistic function of the difference between the item location parameter and the subject location parameter along a single continuous latent dimension. Our results have shown that the Cronbach's alpha was very high (.902) and that the item-total correlations were quite satisfactory (mean .470), thus giving a strong impression of homogeneity (similar to unidimensionality for many authors).(ABSTRACT TRUNCATED AT 250 WORDS)

Aged↗

Modeling and simulation of biological systems with stochasticity.

Mathematical modeling is a powerful approach for understanding the complexity of biological systems. Recently, several successful attempts have been made for simulating complex biological processes like metabolic pathways, gene regulatory networks and cell signaling pathways. The pathway models have not only generated experimentally verifiable hypothesis but have also provided valuable insights into the behavior of complex biological systems. Many recent studies have confirmed the phenotypic variability of organisms to an inherent stochasticity that operates at a basal level of gene expression. Due to this reason, development of novel mathematical representations and simulations algorithms are critical for successful modeling efforts in biological systems. The key is to find a biologically relevant representation for each representation. Although mathematically rigorous and physically consistent, stochastic algorithms are computationally expensive, they have been successfully used to model probabilistic events in the cell. This paper offers an overview of various mathematical and computational approaches for modeling stochastic phenomena in cellular systems.

Algorithms↗

A Markov model for the assembly of heterochromatic regions in position effect variegation.

Here we give a mathematical model for the assembly of heterochromatic regions at the heterochromatin-euchromatin interface in position effect variegation. This probabilistic model predicts the proportions of cells in which a gene is active in cells with one and two variegating chromosomes. The association of heterochromatic proteins to form remodeled chromatin following DNA replication is mainly described by accumulation independent conditional probabilities. These probabilities are conditional on the boundary of the sites to which the proteins can bind; they give the relative attractiveness of the sites to a protein complex chosen at random from a pool of available complexes. The number of complexes available is assumed to be limited and rates of reaction are implicitly modeled by the conditional probabilities. In general, these conditional probabilities are not known, however, they can be experimentally determined. By comparing double variegation situations to single variegation, this model shows that there may be an effect on the expression of reporter genes located near the interfaces due to different sites competing for heterochromatic proteins. In addition, this model suggests that in some cases the attractiveness of sites may change in the presence of other chemical species. Consequently, the model distinguishes between two sorts of data obtained from competition experiments using position effect variegation. The two sorts of data differ as to whether there is a change in the attractiveness of sites in addition to an effect from different sites competing for the same constituents of heterochromatin. Subject to the fact that some of its parameters are not known precisely, this model replicates data from several experiments and can give predictions in other cases.

Animals↗

A tutorial on Markov models based on Mendel's classical experiments.

Hidden Markov Models (HMM) can be extremely useful tools for the analysis of data from biological sequences, and provide a probabilistic model of protein families. Most reviews and general introductions follow the excellent tutorial by Rabiner, where the focus is outside biology. Mendel's famous experiments in plant hybridisation were published in 1866 and are often considered the icebreaking work of modern genetics. He had no prior knowledge of the dual nature of genes, but through a series of experiments he was able to anticipate the hidden concept and name it "Elemente". In this paper we present the background, theory and algorithms of HMM based on examples from Mendel's experiments, and introduce the toolbox "mendelHMM". This approach is considered to have some intuitive advantages in a biological and bioinformatical setting. Applications to analysing bio-sequences like nucleic acids and proteins are also discussed.

Animals↗

Spontaneous regression of residual tumour burden: prediction by Monte Carlo simulation.

Current cancer treatment protocols are designed to release the tumour burden down to a small number of cells. In this study, we use Monte Carlo simulations to show that small populations of cells with intrinsic cell loss rates comparable to the cell loss rates observed clinically in human tumours, may regress spontaneously. Large populations of cells tend to grow under the same conditions of cell loss that result in extinction of small clones. Furthermore, minor variations in the intrinsic cell death probability near 0.50 result in large differences in the number of surviving cells calculated at the 100th generation. When Monte Carlo simulations of clonal growth resulted in clones with large populations (> 50 cells), the population as a whole behaved in a deterministic fashion (logarithmic growth) similar to those observed in clinically observed neoplasms and consistent with other published models of tumour growth. These findings provide a plausible explanation for the clinically observed failure of tumours to recur in instances where tumour burden remains following cancer therapy. The findings also demonstrate the usefulness of the Monte Carlo method to simulate biologic events in populations where the fate of each member of a population can be modeled probabilistically.

Cell Death↗

The future of dynamic systems models in developmental psychology in the light of the past.

A compact glance at the history and impact of mathematical models in development provides background for predicting the fate of dynamical systems modeling in developmental psychology. Dynamic models are considered and the articles by Thelen and Ulrich (1991) and van Geert (1991) are summarized. Deterministic and probabilistic models are compared and some cautions are presented along with a consideration of forms that can be attained by linear models in comparison to those well known in logistic difference models. Reactions to the research by Rabinowitz, Grant, Howe, and Walsh; Kreindler and Lumsden; and Cooney and Troyer are given, with some concluding remarks on survival of dynamical systems modeling in developmental psychology.

Child↗

Likelihood-based refinement. I. Irremovable model errors.

In conventional structure refinement, the discrepancy between the calculated magnitudes and those observed in X-ray experiments is attributed to errors inherent in preliminary assigned values of the model parameters. However, the chosen set of model parameters may not be adequate to describe the structure factors precisely. For example, if some atoms are not included in the current model, then the structure factors calculated from such a partial model contain 'irremovable errors'. These errors cannot be eliminated by any choice of the parameters of the partial structure. Probabilistic modelling suggests a way to take irremovable errors into account. Every trial set of values of the model parameters is now associated with the joint probability distribution of the calculated magnitudes, rather than with a particular set of magnitudes. The new goal of the refinement is formulated as the search for the distribution that is the most consistent with the observed data. The statistical likelihood is a possible measure of the consistency. The suggested quadratic approximation of the likelihood function allows the likelihood-based refinement to be considered as a kind of least-squares refinement that uses appropriate weights and modified targets for the calculated magnitudes. This in turn enables the analysis of tendencies of the likelihood-based refinement in comparison with the classical least-squares refinement.

Journal Article↗

An EM algorithm for dynamic SPECT.

In this paper we present two variants of the EM algorithm for dynamic SPECT imaging. A version based on compartmental modeling which fits a sum of exponentials and a more general approach allowing for arbitrary decaying activities. The underlying probabilistic models are discussed and the incomplete and complete data spaces are shown to be physically meaningful. We indicate that the second method, leading to a convex program in the M step, is easier to treat numerically and we present a possible numerical approach. Some preliminary numerical tests indicating the feasibility of the method are included.

Algorithms↗

An EM algorithm for the block mixture model.

Although many clustering procedures aim to construct an optimal partition of objects or, sometimes, of variables, there are other methods, called block clustering methods, which consider simultaneously the two sets and organize the data into homogeneous blocks. Recently, we have proposed a new mixture model called block mixture model which takes into account this situation. This model allows one to embed simultaneous clustering of objects and variables in a mixture approach. We have studied this probabilistic model under the classification likelihood approach and developed a new algorithm for simultaneous partitioning based on the Classification EM algorithm. In this paper, we consider the block clustering problem under the maximum likelihood approach and the goal of our contribution is to estimate the parameters of this model. Unfortunately, the application of the EM algorithm for the block mixture model cannot be made directly; difficulties arise due to the dependence structure in the model and approximations are required. Using a variational approximation, we propose a generalized EM algorithm to estimate the parameters of the block mixture model and, to illustrate our approach, we study the case of binary data by using a Bernoulli block mixture.

Algorithms↗

Ecological statistics of Gestalt laws for the perceptual organization of contours.

Although numerous studies have measured the strength of visual grouping cues for controlled psychophysical stimuli, little is known about the statistical utility of these various cues for natural images. In this study, we conducted experiments in which human participants trace perceived contours in natural images. These contours are automatically mapped to sequences of discrete tangent elements detected in the image. By examining relational properties between pairs of successive tangents on these traced curves, and between randomly selected pairs of tangents, we are able to estimate the likelihood distributions required to construct an optimal Bayesian model for contour grouping. We employed this novel methodology to investigate the inferential power of three classical Gestalt cues for contour grouping: proximity, good continuation, and luminance similarity. The study yielded a number of important results: (1) these cues, when appropriately defined, are approximately uncorrelated, suggesting a simple factorial model for statistical inference; (2) moderate image-to-image variation of the statistics indicates the utility of general probabilistic models for perceptual organization; (3) these cues differ greatly in their inferential power, proximity being by far the most powerful; and (4) statistical modeling of the proximity cue indicates a scale-invariant power law in close agreement with prior psychophysics.

Form Perception↗

Probabilistic gas and bubble dynamics models of decompression sickness occurrence in air and nitrogen-oxygen diving.

Probabilistic models of the occurrence of decompression sickness (DCS) with instantaneous risk defined as the weighted sum of bubble volumes in each of three parallel-perfused gas exchange compartments were fit using likelihood maximization to the subset of the USN Primary Air and N2-O2 database [n = 2,383, mean P(DCS) = 5.8%] used in development of the USN LE1 probabilistic models. Bubble dynamics with one diffusible gas in each compartment were modeled using the Van Liew equations with the nucleonic bubble radius, compartmental volume, compartmental bulk N2 diffusivity, compartmental N2 solubility, and the N2 solubility in blood x compartmental blood flow as adjustable parameters. Models were also tested that included the effects of linear elastic resistance to bubble growth in one, two, or all three of the modeled compartments. Model performance about the training data and separate validation data was compared to results obtained about the same data using the LE1 probabilistic model, which was independently implemented from published descriptions. In the most successful bubble volume model, BVM(3), diffusion significantly slows bubble growth in one of the modeled compartments, whereas mechanical resistance to bubble growth substantially accelerates bubble resolution in all compartments. BVM(3) performed generally on a par with LE1, despite inclusion of 12 more adjustable parameters, and tended to provide more accurate incidence-only estimates of DCS probability than LE1, particularly for profiles in which high fractional O2 gas mixes are breathed. Values of many estimated BVM(3) parameters were outside of the physiologic range, indicating that the model emerged from optimization as a mathematical descriptor of processes beyond bubble formation and growth that also contribute to DCS outcomes. Although incomplete as a mechanistic description of DCS etiology, BVM(3) remains applicable to a wider variety of decompressions than LE1 and affords a conceptual framework for further refinements motivated by mechanistic principles.

Decompression Sickness↗

Context-specific Bayesian clustering for gene expression data.

The recent growth in genomic data and measurements of genome-wide expression patterns allows us to apply computational tools to examine gene regulation by transcription factors. In this work, we present a class of mathematical models that help in understanding the connections between transcription factors and functional classes of genes based on genetic and genomic data. Such a model represents the joint distribution of transcription factor binding sites and of expression levels of a gene in a unified probabilistic model. Learning a combined probability model of binding sites and expression patterns enables us to improve the clustering of the genes based on the discovery of putative binding sites and to detect which binding sites and experiments best characterize a cluster. To learn such models from data, we introduce a new search method that rapidly learns a model according to a Bayesian score. We evaluate our method on synthetic data as well as on real life data and analyze the biological insights it provides. Finally, we demonstrate the applicability of the method to other data analysis problems in gene expression data.

Bayes Theorem↗

Inferring quantitative models of regulatory networks from expression data.

MOTIVATION: Genetic networks regulate key processes in living cells. Various methods have been suggested to reconstruct network architecture from gene expression data. However, most approaches are based on qualitative models that provide only rough approximations of the underlying events, and lack the quantitative aspects that are critical for understanding the proper function of biomolecular systems. RESULTS: We present fine-grained dynamical models of gene transcription and develop methods for reconstructing them from gene expression data within the framework of a generative probabilistic model. Unlike previous works, we employ quantitative transcription rates, and simultaneously estimate both the kinetic parameters that govern these rates, and the activity levels of unobserved regulators that control them. We apply our approach to expression datasets from yeast and show that we can learn the unknown regulator activity profiles, as well as the binding affinity parameters. We also introduce a novel structure learning algorithm, and demonstrate its power to accurately reconstruct the regulatory network from those datasets.

Binding Sites↗

A multiple-feature framework for modelling and predicting transcription factor binding sites.

MOTIVATION: The identification of transcription factor binding sites in promoter sequences is an important problem, since it reveals information about the transcriptional regulation of genes. For analysing transcriptional regulation, computational approaches for predicting putative binding sites are applied. Commonly used stochastic models for binding sites are position-specific score matrices, which show weak predictive power. RESULTS: We have developed a probabilistic modelling approach, which allows to consider diverse characteristic binding site properties to obtain more accurate representations of binding sites. These properties are modelled as random variables in Bayesian networks, which are capable of dealing with dependencies among binding site properties. Cross-validation on several datasets shows improvements in the false positive error rate and the significance (P-value) of true binding sites.

Algorithms↗

Sensitivity analysis of a two-dimensional probabilistic risk assessment model using analysis of variance.

This article demonstrates application of sensitivity analysis to risk assessment models with two-dimensional probabilistic frameworks that distinguish between variability and uncertainty. A microbial food safety process risk (MFSPR) model is used as a test bed. The process of identifying key controllable inputs and key sources of uncertainty using sensitivity analysis is challenged by typical characteristics of MFSPR models such as nonlinearity, thresholds, interactions, and categorical inputs. Among many available sensitivity analysis methods, analysis of variance (ANOVA) is evaluated in comparison to commonly used methods based on correlation coefficients. In a two-dimensional risk model, the identification of key controllable inputs that can be prioritized with respect to risk management is confounded by uncertainty. However, as shown here, ANOVA provided robust insights regarding controllable inputs most likely to lead to effective risk reduction despite uncertainty. ANOVA appropriately selected the top six important inputs, while correlation-based methods provided misleading insights. Bootstrap simulation is used to quantify uncertainty in ranks of inputs due to sampling error. For the selected sample size, differences in F values of 60% or more were associated with clear differences in rank order between inputs. Sensitivity analysis results identified inputs related to the storage of ground beef servings at home as the most important. Risk management recommendations are suggested in the form of a consumer advisory for better handling and storage practices.

Analysis of Variance↗

Cross-species comparison significantly improves genome-wide prediction of cis-regulatory modules in Drosophila.

BACKGROUND: The discovery of cis-regulatory modules in metazoan genomes is crucial for understanding the connection between genes and organism diversity. It is important to quantify how comparative genomics can improve computational detection of such modules. RESULTS: We run the Stubb software on the entire D. melanogaster genome, to obtain predictions of modules involved in segmentation of the embryo. Stubb uses a probabilistic model to score sequences for clustering of transcription factor binding sites, and can exploit multiple species data within the same probabilistic framework. The predictions are evaluated using publicly available gene expression data for thousands of genes, after careful manual annotation. We demonstrate that the use of a second genome (D. pseudoobscura) for cross-species comparison significantly improves the prediction accuracy of Stubb, and is a more sensitive approach than intersecting the results of separate runs over the two genomes. The entire list of predictions is made available online. CONCLUSION: Evolutionary conservation of modules serves as a filter to improve their detection in silico. The future availability of additional fruitfly genomes therefore carries the prospect of highly specific genome-wide predictions using Stubb.

Algorithms↗

Dietary exposure to chemical migrants from food contact materials: a probabilistic approach.

A two-dimensional probabilistic model has been developed to estimate the short-term dietary exposure of UK consumers to migrants from food packaging materials. The current EU approach uses a default scenario of assuming that all individuals are 60 kg weight and consume 1 kg of food packaged in the material of interest per day. Using four UK National Dietary and Nutrition Surveys comprising 4-7 day dietary records for different age groups and survey years, a sample representative of the UK population has been obtained consuming around 4200 different food items. Each survey provides records for around 2000 individuals and supplies detailed information on the consumption of food and data on sex, height and socio-economic status which may be used to analyse the exposure of selected groups within the community. As a result we are able to address the variation in consumption of food amongst individuals, and account for actual body weights providing a more accurate representation of the 'true' exposure. The migrants bisphenol A diglycidyl ether (BADGE), di-2-ethylhexyl adipate (DEHA) and styrene were considered as specimen compounds although the methodology employed has the flexibility to adapt to other migrants and packaging types and indeed other food contaminants. Exposure for each individual is estimated by calculating and summing the individual exposure from each item in their diet, and is repeated for all individuals in each survey to produce a distribution of exposures for the population. The packaging type of each food item is assigned by utilizing known packaging types from the database or, by sampling from a distribution based upon market share information. The parameters contributing towards the exposure from a packaged dietary item are migrant concentration and item weight. Distributions are used to represent the inherent variation and uncertainty affecting these parameters. Where data on concentrations for a particular type of food are lacking, expert judgement is used to extrapolate from available data for other food types. The model can also be run using only migration data for food simulants. In this case, concentrations expected for each of the food items are assigned based on the data for the relevant food simulant. The primary outputs of the model are distributions of estimated daily intakes for the selected population. Each distribution gives the variation across the population subject to the uncertain parameters sampled in that iteration of the model. Analysing the ensemble of distributions allows us to obtain the confidence limits around estimates for percentiles due to the uncertainties. The probabilistic approach allows sensitivity analysis to evaluate the relative importance of the input parameters and places confidence bounds on the outputs to show the effect of the uncertainties and the contribution of each food type toward the overall exposure.

Adipates↗