Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

A discriminative model for identifying spatial cis-regulatory modules.

Transcriptional regulation is mediated by the coordinated binding of transcription factors to the upstream regions of genes. In higher eukaryotes, the binding sites of cooperating transcription factors are organized into short sequence units, called cis-regulatory modules. In this paper, we propose a method for identifying modules of transcription factor binding sites in a set of co-regulated genes, using only the raw sequence data as input. Our method is based on a novel probabilistic model that describes the mechanism of cis-regulation, including the binding sites of cooperating transcription factors, the organization of these binding sites into short sequence modules, and the regulation of a gene by its modules. We show that our method is successful in discovering planted modules in simulated data and known modules in yeast. More importantly, we applied our method to a large collection of human gene sets and found 83 significant cis-regulatory modules, which included 36 known motifs and many novel ones. Thus, our results provide one of the first comprehensive compendiums of putative cis-regulatory modules in human.

Binding Sites↗

A model-based approach to selection of tag SNPs.

BACKGROUND: Single Nucleotide Polymorphisms (SNPs) are the most common type of polymorphisms found in the human genome. Effective genetic association studies require the identification of sets of tag SNPs that capture as much haplotype information as possible. Tag SNP selection is analogous to the problem of data compression in information theory. According to Shannon's framework, the optimal tag set maximizes the entropy of the tag SNPs subject to constraints on the number of SNPs. This approach requires an appropriate probabilistic model. Compared to simple measures of Linkage Disequilibrium (LD), a good model of haplotype sequences can more accurately account for LD structure. It also provides a machinery for the prediction of tagged SNPs and thereby to assess the performances of tag sets through their ability to predict larger SNP sets. RESULTS: Here, we compute the description code-lengths of SNP data for an array of models and we develop tag SNP selection methods based on these models and the strategy of entropy maximization. Using data sets from the HapMap and ENCODE projects, we show that the hidden Markov model introduced by Li and Stephens outperforms the other models in several aspects: description code-length of SNP data, information content of tag sets, and prediction of tagged SNPs. This is the first use of this model in the context of tag SNP selection. CONCLUSION: Our study provides strong evidence that the tag sets selected by our best method, based on Li and Stephens model, outperform those chosen by several existing methods. The results also suggest that information content evaluated with a good model is more sensitive for assessing the quality of a tagging set than the correct prediction rate of tagged SNPs. Besides, we show that haplotype phase uncertainty has an almost negligible impact on the ability of good tag sets to predict tagged SNPs. This justifies the selection of tag SNPs on the basis of haplotype informativeness, although genotyping studies do not directly assess haplotypes. A software that implements our approach is available.

Databases, Genetic↗

A fast nonlinear method for parametric imaging of myocardial perfusion by dynamic (13)N-ammonia PET.

UNLABELLED: A parametric image of myocardial perfusion (mL/min/g) is a quantitative image generated by fitting a tracer kinetic model to dynamic (13)N-ammonia PET data on a pixel-by-pixel basis. There are several methods for such parameter estimation problems, including weighted nonlinear regression (WNLR) and a fast linearizing method known as Patlak analysis. Previous work showed that sigmoidal networks can be used for parameter estimation of mono- and biexponential models. The method used in this study is a hybrid of WNLR and sigmoidal networks called nonlinear regression estimation (NRE). The purpose of the study is to compare NRE with WNLR and Patlak analysis for parametric imaging of perfusion in the canine heart by (13)N-ammonia PET. METHODS: A simulation study measured the statistical performance of NRE, WNLR, and Patlak analysis for a probabilistic model of time-activity curves. Four canine subjects were injected with 740 MBq (13)N-ammonia and scanned dynamically. Images were reconstructed with filtered backprojection and resliced into short-axis cuts. Parametric images of a single midventricular plane per subject were generated by NRE, WNLR, and Patlak analysis. Small regions of interest (ROIs) were drawn on each parametric image (8 ROIs per subject for a total of 32). RESULTS: For the simulation study, the median absolute value of the relative error for a perfusion value of 1.0 mL/min/g was 16.6% for NRE, 17.9% for WNLR, 19.5% for Patlak analysis, and 14.5% for an optimal WNLR method (computable by simulation only). All methods are unbiased conditioned on a wide range of perfusion values. For the canine studies, the least squares line fits comparing NRE (y) and Patlak analysis (z) with WNLR (x) for all 32 ROIs were y = 1.02x - 0.028 and z = 0.90x + 0.019, respectively. Both NRE and Patlak analysis generate 128 x 128 parametric images in seconds. CONCLUSION: The statistical performance of NRE is competitive with WNLR and superior to Patlak analysis for parametric imaging of myocardial perfusion. NRE is a fast nonlinear alternative to Patlak analysis and other fast linearizing methods for parametric imaging. NRE should be applicable to many other tracers and tracer kinetic models.

Ammonia↗

Do all daughter cells enter the "indeterminate" ("A") state of the cell cycle? Analysis of stathmokinetic experiments on L1210 cells.

The results of ten different stathmokinetic experiments on L1210 cells are analyzed to determine whether all the postmitotic cells enter the exponential AG1 compartment of the cell cycle characterized by exponentially distributed transit times. The analysis is based on a mathematical model coherent with the generalized A-B-transition hypothesis. An original fitting procedure is introduced to estimate the fraction of cells entering AG1 as well as other parameters of the cell cycle. Results of the analysis suggest that either all or nearly all postmitotic cells enter G1A. The results are discussed with respect to the validity of the A-B-transition hypothesis of the "probabilistic" model of the cell cycle. Some systematically occurring discrepancies that do not conform with generally accepted cell cycle models are apparent from the analysis of this data.

Animals↗

Recovering slant and angular velocity from a linear velocity field: modeling and psychophysics.

The data from two experiments, both using stimuli simulating orthographically rotating surfaces, are presented, with the primary variable of interest being whether the magnitude of the simulated gradient was from expanding vs. contracting motion. One experiment asked observers to report the apparent slant of the rotating surface, using a gauge figure. The other experiment asked observers to report the angular velocity, using a comparison rotating sphere. The results from both experiments clearly show that observers are less sensitive to expanding than to contracting optic-flow fields. These results are well predicted by a probabilistic model which derives the orientation and angular velocity of the projected surface from the properties of the optic flow computed within an extended time window.

Depth Perception↗

Spike sorting: Bayesian clustering of non-stationary data.

Spike sorting involves clustering spikes recorded by a micro-electrode according to the source neurons. It is a complicated task, which requires much human labor, in part due to the non-stationary nature of the data. We propose to automate the clustering process in a Bayesian framework, with the source neurons modeled as a non-stationary mixture-of-Gaussians. At a first search stage, the data are divided into short time frames, and candidate descriptions of the data as mixtures-of-Gaussians are computed for each frame separately. At a second stage, transition probabilities between candidate mixtures are computed, and a globally optimal clustering solution is found as the maximum-a-posteriori solution of the resulting probabilistic model. The transition probabilities are computed using local stationarity assumptions, and are based on a Gaussian version of the Jensen-Shannon divergence. We employ synthetically generated spike data to illustrate the method and show that it outperforms other spike sorting methods in a non-stationary scenario. We then use real spike data and find high agreement of the method with expert human sorters in two modes of operation: a fully unsupervised and a semi-supervised mode. Thus, this method differs from other methods in two aspects: its ability to account for non-stationary data, and its close to human performance.

Algorithms↗

An Approval-Voting Polytope for Linear Orders

A probabilistic model of approval voting on n alternatives generates a collection of probability distributions on the family of all subsets of the set of alternatives. Focusing on the size-independent model proposed by Falmagne and Regenwetter, we recast the problem of characterizing these distributions as the search for a minimal system of linear equations and inequalities for a specific convex polytope. This approval-voting polytope, with n! vertices in a space of dimension 2(n), is proved to be of dimension 2(n)-n-1. Several families of facet-defining linear inequalities are exhibited, each of which has a probabilistic interpretation. Some proofs rely on special sequences of rankings of the alternatives. Although the equations and facet-defining inequalities found so far yield a complete minimal description when n<=4 (as indicated by the PORTA software), the problem remains open for larger values of n.

Journal Article↗

A graph-based motif detection algorithm models complex nucleotide dependencies in transcription factor binding sites.

Given a set of known binding sites for a specific transcription factor, it is possible to build a model of the transcription factor binding site, usually called a motif model, and use this model to search for other sites that bind the same transcription factor. Typically, this search is performed using a position-specific scoring matrix (PSSM), also known as a position weight matrix. In this paper we analyze a set of eukaryotic transcription factor binding sites and show that there is extensive clustering of similar k-mers in eukaryotic motifs, owing to both functional and evolutionary constraints. The apparent limitations of probabilistic models in representing complex nucleotide dependencies lead us to a graph-based representation of motifs. When deciding whether a candidate k-mer is part of a motif or not, we base our decision not on how well the k-mer conforms to a model of the motif as a whole, but how similar it is to specific, known k-mers in the motif. We elucidate the reasons why we expect graph-based methods to perform well on motif data. Our MotifScan algorithm shows greatly improved performance over the prevalent PSSM-based method for the detection of eukaryotic motifs.

Algorithms↗

One-shot learning of object categories.

Learning visual models of object categories notoriously requires hundreds or thousands of training examples. We show that it is possible to learn much information about a category from just one, or a handful, of images. The key insight is that, rather than learning from scratch, one can take advantage of knowledge coming from previously learned categories, no matter how different these categories might be. We explore a Bayesian implementation of this idea. Object categories are represented by probabilistic models. Prior knowledge is represented as a probability density function on the parameters of these models. The posterior model for an object category is obtained by updating the prior in the light of one or more observations. We test a simple implementation of our algorithm on a database of 101 diverse object categories. We compare category models learned by an implementation of our Bayesian approach to models learned from by Maximum Likelihood (ML) and Maximum A Posteriori (MAP) methods. We find that on a database of more than 100 categories, the Bayesian approach produces informative models when the number of training examples is too small for other methods to operate successfully.

Algorithms↗

Bayesian fMRI data analysis with sparse spatial basis function priors.

In previous work we have described a spatially regularised General Linear Model (GLM) for the analysis of brain functional Magnetic Resonance Imaging (fMRI) data where Posterior Probability Maps (PPMs) are used to characterise regionally specific effects. The spatial regularisation is defined over regression coefficients via a Laplacian kernel matrix and embodies prior knowledge that evoked responses are spatially contiguous and locally homogeneous. In this paper we propose to finesse this Bayesian framework by specifying spatial priors using Sparse Spatial Basis Functions (SSBFs). These are defined via a hierarchical probabilistic model which, when inverted, automatically selects an appropriate subset of basis functions. The method includes non-linear wavelet shrinkage as a special case. As compared to Laplacian spatial priors, SSBFs allow for spatial variations in signal smoothness, are more computationally efficient and are robust to heteroscedastic noise. Results are shown on synthetic data and on data from an event-related fMRI experiment.

Algorithms↗

Modeling the signaling endosome hypothesis: why a drive to the nucleus is better than a (random) walk.

BACKGROUND: Information transfer from the plasma membrane to the nucleus is a universal cell biological property. Such information is generally encoded in the form of post-translationally modified protein messengers. Textbook signaling models typically depend upon the diffusion of molecular signals from the site of initiation at the plasma membrane to the site of effector function within the nucleus. However, such models fail to consider several critical constraints placed upon diffusion by the cellular milieu, including the likelihood of signal termination by dephosphorylation. In contrast, signaling associated with retrogradely transported membrane-bounded organelles such as endosomes provides a dephosphorylation-resistant mechanism for the vectorial transmission of molecular signals. We explore the relative efficiencies of signal diffusion versus retrograde transport of signaling endosomes. RESULTS: Using large-scale Monte Carlo simulations of diffusing STAT-3 molecules coupled with probabilistic modeling of dephosphorylation kinetics we found that predicted theoretical measures of STAT-3 diffusion likely overestimate the effective range of this signal. Compared to the inherently nucleus-directed movement of retrogradely transported signaling endosomes, diffusion of STAT-3 becomes less efficient at information transfer in spatial domains greater than 200 nanometers from the plasma membrane. CONCLUSION: Our model suggests that cells might utilize two distinct information transmission paradigms: 1) fast local signaling via diffusion over spatial domains on the order of less than 200 nanometers; 2) long-distance signaling via information packets associated with the cytoskeletal transport apparatus. Our model supports previous observations suggesting that the signaling endosome hypothesis is a subset of a more general hypothesis that the most efficient mechanism for intracellular signaling-at-a-distance involves the association of signaling molecules with molecular motors that move along the cytoskeleton. Importantly, however, cytoskeletal association of membrane-bounded complexes containing ligand-occupied transmembrane receptors and downstream effector molecules provides the ability to regenerate signals at any point along the transmission path. We conclude that signaling endosomes provide unique information transmission properties relevant to all cell architectures, and we propose that the majority of relevant information transmitted from the plasma membrane to the nucleus will be found in association with organelles of endocytic origin.

Animals↗

Model dependence in quantification of spike interdependence by joint peri-stimulus time histogram.

Multineuronal recordings have enabled us to examine context-dependent changes in the relationship between the activities of multiple cells. The joint peri-stimulus time histogram (JPSTH) is a much-used method for investigating the dynamics of the interdependence of spike events between pairs of cells. Its results are often taken as an estimate of interaction strength between cells, independent of modulations in the cells' firing rates. We evaluate the adequacy of this estimate by examining the mathematical structure of how the JPSTH quantifies an interaction strength after excluding the contribution of firing rates. We introduce a simple probabilistic model of interacting point processes to generate simulated spike data and show that the normalized JPSTH incorrectly infers the temporal structure of variations in the interaction parameter strength. This occurs because, in our model, the correct normalization of firing-rate contributions is different than that used in Aertsen, Gerstein, Habib, and Palm's (1989) effective connectivity model. This demonstrates that firing-rate modulations cannot be corrected for in a model-independent manner, and therefore the effective connectivity does not represent a universal characteristic that is independent of modulation of the firing rates. Aertsen et al.'s (1989) effective connectivity may still be used in the analysis of experimental data, provided we are aware that this is simply one of many ways of describing the structure of interdependence. We also discuss some measure-independent characteristics of the structure of interdependence.

Brain↗

The utility of extended longitudinal profiles in predicting future health care expenditures.

BACKGROUND: Health care spending is highly concentrated. Prediction models that accurately identify the characteristics of individuals most likely to incur high levels of health expenditures in a subsequent year are important analytical and statistical tools. OBJECTIVES: This study examined the capacity of alternative models to predict the likelihood of incurring high levels of medical expenditures in a subsequent year. This effort also evaluated the utility of an additional year of longitudinal information. SUBJECTS: A nationally representative sample from the Medical Expenditure Panel Survey (MEPS). METHODS: The MEPS longitudinal data are used to examine the persistence of high expenditures during a 2-year period. With the unique linkage of the MEPS to the National Health Interview Survey, the utility of an additional year of data also was examined. Resultant models were evaluated in terms of sensitivity, specificity, and predictive capacity. RESULTS: Only modest marginal gains in discrimination capacity were realized from the use of extended longitudinal profiles from the National Health Interview Survey, relative to information on prior year characteristics. CONCLUSIONS: Our results highlight the continuing concentration of health care expenditures during the period 1996 to 2002 and reveal some attenuation in magnitude in the tail of this distribution over time. Further, our results provide evidence of the utility of probabilistic models as prediction tools to identify individuals likely to incur high levels of expenditures in future years. Predictive capacity does not suffer when restricted to a single year of prior information.

Adolescent↗

Assessing the risks of pesticide residues to consumers: recent and future developments.

Assessing exposure of consumers to pesticide residues is an area of regulatory science that has rapidly developed over the last decade. From simplistic, deterministic models calculating lifetime exposure for adults only, assessment procedures have diversified so that more realistic estimates of long term exposures for adults, schoolchildren, toddlers and infants and short term exposures for adults and toddlers (who generally bound the more extreme consumer patterns) are now carried out. The final assessment of risk still remains a simplistic numeric comparison against hazard assessment based on a wide range of toxicity studies incorporating the appropriate safety or uncertainty factors. As development of risk assessments continues, the use of probabilistic models is becoming an invaluable information tool for quantitative risk management and aiding assessment of cumulative exposure. This paper examines the recent developments in risk assessment and consumer perception of the risks of pesticide residues, and speculates where the future developments in these areas may lie.

Adolescent↗

Realistic protein-protein association rates from a simple diffusional model neglecting long-range interactions, free energy barriers, and landscape ruggedness.

We develop a simple but rigorous model of protein-protein association kinetics based on diffusional association on free energy landscapes obtained by sampling configurations within and surrounding the native complex binding funnels. Guided by results obtained on exactly solvable model problems, we transform the problem of diffusion in a potential into free diffusion in the presence of an absorbing zone spanning the entrance to the binding funnel. The free diffusion problem is solved using a recently derived analytic expression for the rate of association of asymmetrically oriented molecules. Despite the required high steric specificity and the absence of long-range attractive interactions, the computed rates are typically on the order of 10(4)-10(6) M(-1) sec(-1), several orders of magnitude higher than rates obtained using a purely probabilistic model in which the association rate for free diffusion of uniformly reactive molecules is multiplied by the probability of a correct alignment of the two partners in a random collision. As the association rates of many protein-protein complexes are also in the 10(5)-10(6) M(-1) sec(-1) range, our results suggest that free energy barriers arising from desolvation and/or side-chain freezing during complex formation or increased ruggedness within the binding funnel, which are completely neglected in our simple diffusional model, do not contribute significantly to the dynamics of protein-protein association. The transparent physical interpretation of our approach that computes association rates directly from the size and geometry of protein-protein binding funnels makes it a useful complement to Brownian dynamics simulations.

Computer Simulation↗

Modeling the percolation of annotation errors in a database of protein sequences.

Public sequence databases contain information on the sequence, structure and function of proteins. Genome sequencing projects have led to a rapid increase in protein sequence information, but reliable, experimentally verified, information on protein function lags a long way behind. To address this deficit, functional annotation in protein databases is often inferred by sequence similarity to homologous, annotated proteins, with the attendant possibility of error. Now, the functional annotation in these homologous proteins may itself have been acquired through sequence similarity to yet other proteins, and it is generally not possible to determine how the functional annotation of any given protein has been acquired. Thus the possibility of chains of misannotation arises, a process we term 'error percolation'. With some simple assumptions, we develop a dynamical probabilistic model for these misannotation chains. By exploring the consequences of the model for annotation quality it is evident that this iterative approach leads to a systematic deterioration of database quality.

Animals↗

Graphical and statistical approaches to data analysis for in situ hybridization.

Quantification of gene expression in a morphological context is an invaluable tool for neurobiological investigation. The ability to measure the quantity of specific mRNA molecules at the level of the single neuron permits one to monitor the modulation of complex cell synthetic activity of intact neuron populations. The cells of interest can be contiguous or dispersed in functionally significant patterns throughout a broad anatomical region of the brain. The application of quantitative in situ hybridization is technically difficult and labor intensive. Nevertheless, it has great utility for investigating gene expression from a structural perspective. (1) In situ hybridization permits one to ask questions concerning the anatomical pattern of neuronal gene expression. (2) It permits analyses concerning the initiation of expression, cell location, cell type, and alterations of level of expression within a spatial and temporal context. (3) In cases where blotting methods suggest a message exists at low copy, in situ hybridization permits queries at the single-cell level. For example, in situ hybridization can determine if very few cells are expressing the gene product or if many neurons dispersed throughout a brain region exhibit low mRNA copy number/cell. Quantitative analyses also allow detailed investigation of cell response to physiologically meaningful stimulation. Our application of statistical and numerical methods is a demonstration of the utility of probabilistic models; the mixture distribution accounted for data from both labeled and unlabeled sources. In agreement with many previous investigations, grain density over an unlabeled uniform source (oxytocinergic cells) was suitably described by the Poisson distribution. The population of labeled vasopressinergic cells, however, was best described by the negative binomial distribution. Previous investigations from different fields of biology show that the negative binomial can be used to describe many biological phenomena, and this distribution was considered in at least two previous investigations to evaluate autoradiographic data which did not fit the Poisson function. From a theoretical perspective, the probabilistic relationship between beta-particle decay (a Poisson function) and the distribution of message levels among individual neurons in a cell group (gamma distribution) prompts consideration of the negative binomial. For both data sets the observed variances were larger than the mean, and the labeled portion of the data sets exhibited positive skewness.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

Monte Carlo simulation of hyaluronidase reaction involving hydrolysis, transglycosylation and condensation.

The action of hyaluronidase on oligosaccharides from hyaluronan is complicated due to branched reaction paths containing hydrolysis, transglycosylation and condensation. The unit component of hyaluronan is a disaccharide, namely GlcA-(beta 1-->3)-GlcNAc where GlcA and GlcNAc are d-glucuronic acid and d-N-acetylglucosamine respectively. Hyaluronan is the linear polymer formed by these disaccharide units, linked together with beta 1-->4 glycosidic bonds. Bovine testicular hyaluronidase acts only at beta 1-->4 glycosidic bonds of hyaluronan. The progress of product distribution from short oligosaccharides was simulated with the Monte Carlo method using the probabilistic model. The model consists only of a single enzyme molecule and a finite number of substrate and water molecules. The simulation is based on a simple reaction scheme and proceeds via an algorithm with minimum adjustable parameters generating random numbers and probabilities. The experimental data for bovine testicular hyaluronidase using [GlcA-(beta 1-->3)-GlcNAc](4) as the starting substrate were quantitatively simulated with only three adjustable parameters. The simulated data for [GlcA-(beta 1-->3)-GlcNAc](3) and [GlcA-(beta 1-->3)-GlcNAc](5) as the starting substrates agreed semi-quantitatively with experimental data using the same parameters. The mechanism of the hyaluronidase reaction is a combination of branched probabilistic cycles. The condensation reaction is much weaker than the transglycosylation reaction but contributes to product distribution at the final stage of the reaction, preventing complete hydrolysis of the substrates.

Algorithms↗