Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Preventing second-generation infections in a smallpox bioterror attack.

This article presents a new probabilistic model for the prevention of second-generation infections by different vaccination strategies in the event of a smallpox bioterror attack. The main results are independent of the reproductive number R0 (the number of secondary infections transmitted per index infected individual) and population mixing patterns. General expressions are derived for the fraction of second-generation infections that can be prevented through vaccination, whereas specific results are obtained for traced and mass vaccination, respectively. Expressions for total outbreak size in controlled epidemics are also presented. The analysis highlights the importance of vaccination logistics in addition to beliefs and assumptions regarding smallpox epidemiology in evaluating alternative responses to a smallpox bioterror attack.

Bioterrorism↗

Regression-based sampling for persons with high health expenditures: evaluating accuracy and yield with the 1997 MEPS.

BACKGROUND: Given the high concentration of health care expenditures among a relatively small percentage of the population, the 1997 Medical Expenditure Panel Survey was designed to learn more about these high expenditure individuals by oversampling them. OBJECTIVE: Oversampling high expenditure individuals enables more precise estimation of what the nation's health care dollar buys and who pays it. It also enhances the ability to discern the causes of high health care expenses and the characteristics of the individuals who incur them. METHOD: Using the 1987 National Medical Expenditure Survey, a probabilistic model was developed to select households from the 1996 National Health Interview Survey likely to contain individuals incurring high levels of medical expenditures in the 1997 MEPS. The accuracy of the selection model, and the degree to which the high expenditure population was oversampled, are assessed with the 1997 MEPS data. RESULTS: Over half of the persons selected by the regression model were expected to have high health expenditures. Of the 456 persons selected by the model for oversampling, 257 individuals or 56.4% did, in fact, have high expenditures. Regression-based sampling increased the proportion of MEPS individuals with high expenditures from 14.3% without oversampling to 17.2% of the total cohort with oversampling (or from 938-1,126 persons). CONCLUSION: This paper demonstrates that a model-based approach to oversampling a high expenditure population, or any population with dynamic characteristics, can be highly successful in terms of sampling yield and accuracy.

Adult↗

The economics of abrupt climate change.

The US National Research Council defines abrupt climate change as a change of state that is sufficiently rapid and sufficiently widespread in its effects that economies are unprepared or incapable of adapting. This may be too restrictive a definition, but abrupt climate change does have implications for the choice between the main response options: mitigation (which reduces the risks of climate change) and adaptation (which reduces the costs of climate change). The paper argues that by (i) increasing the costs of change and the potential growth of consumption, and (ii) reducing the time to change, abrupt climate change favours mitigation over adaptation. Furthermore, because the implications of change are fundamentally uncertain and potentially very high, it favours a precautionary approach in which mitigation buys time for learning. Adaptation-oriented decision tools, such as scenario planning, are inappropriate in these circumstances. Hence learning implies the use of probabilistic models that include socioeconomic feedbacks.

Adaptation, Psychological↗

Content-based indexing of images and video.

By representing image content using probabilistic models of an object's appearance we can obtain semantics-preserving compression of the image data. Such compact representations of an image's salient features allow rapid computer searches of even large image databases. Examples are shown for databases of face images, a video of American sign language (ASL), and a video of facial expressions.

Algorithms↗

A comparative genomics approach to prediction of new members of regulons.

Identifying the complete transcriptional regulatory network for an organism is a major challenge. For each regulatory protein, we want to know all the genes it regulates, that is, its regulon. Examples of known binding sites can be used to estimate the binding specificity of the protein and to predict other binding sites. However, binding site predictions can be unreliable because determining the true specificity of the protein is difficult because of the considerable variability of binding sites. Because regulatory systems tend to be conserved through evolution, we can use comparisons between species to increase the reliability of binding site predictions. In this article, an approach is presented to evaluate the computational predictions of regulatory sites. We combine the prediction of transcription units having orthologous genes with the prediction of transcription factor binding sites based on probabilistic models. We augment the sets of genes in Escherichia coli that are expected to be regulated by two transcription factors, the cAMP receptor protein and the fumarate and nitrate reduction regulatory protein, through a comparison with the Haemophilus influenzae genome. At the same time, we learned more about the regulatory networks of H. influenzae, a species with much less experimental knowledge than E. coli. By studying orthologous genes subject to regulation by the same transcription factor, we also gained understanding of the evolution of the entire regulatory systems.

Amino Acid Sequence↗

Probabilistic description of traffic breakdowns.

We analyze the characteristic features of traffic breakdown. To describe this phenomenon we apply the probabilistic model regarding the jam emergence as the formation of a large car cluster on a highway. In these terms, the breakdown occurs through the formation of a certain critical nucleus in the metastable vehicle flow, which enables us to confine ourselves to one cluster model. We assume that, first, the growth of the car cluster is governed by attachment of cars to the cluster whose rate is mainly determined by the mean headway distance between the car in the vehicle flow and, maybe, also by the headway distance in the cluster. Second, the cluster dissolution is determined by the car escape from the cluster whose rate depends on the cluster size directly. The latter is justified using the available experimental data for the correlation properties of the synchronized mode. We write the appropriate master equation converted then into the Fokker-Planck equation for the cluster distribution function and analyze the formation of the critical car cluster due to the climb over a certain potential barrier. The further cluster growth irreversibly causes jam formation. Numerical estimates of the obtained characteristics and the experimental data of the traffic breakdown are compared. In particular, we draw a conclusion that the characteristic intrinsic time scale of the breakdown phenomenon should be about 1 min and explain the case why the traffic volume interval inside which traffic breakdown is observed is sufficiently wide.

Journal Article↗

Interplay of entropic and memory effects in diffusion of methane in silicalite zeolites.

The role of entropic effects in methane distribution and transport in silicalite zeolites is studied using molecular dynamics in the limit of infinite dilution or small loading. Diffusive behavior and its anisotropy is assessed as a function of temperature where we find both an Arrhenius regime above 250 K and deviations thereof below such temperature. Using a previous probabilistic model, geometrical correlations or memory effects are evidenced and are shown to be enhanced as temperature is reduced. Deviations from Arrhenius behavior are concomitant with entropic effects. We find that, the preference of methane towards presence at intersections or channel centers changes at a threshold temperature. A discrete transition is found from a channel-center preferred phase, at low temperatures, versus an intersection preferred phase at high temperatures with evidence of hysteresis effects. Such entropic effects are also reflected, in diffusive transport, as non-Arrhenius-type behavior. A model based on accessible volume as a function of energy agrees with the simulated transition lending new insight into zeolite cavity design.

Journal Article↗

Bayesian methods for the conformational classification of eight-membered rings.

Two methods for the classification of eight-membered rings based on a Bayesian analysis are presented. The two methods share the same probabilistic model for the measurement of torsion angles, but while the first method uses the canonical forms of cyclooctane and, given an empirical sequence of eight torsion angles, yields the probability that the associated structure corresponds to each of the ten canonical conformations, the second method does not assume previous knowledge of existing conformations and yields a clustering classification of a data set, allowing new conformations to be detected. Both methods have been tested using the conformational classification of Csp3 eight-membered rings described in the literature. The methods have also been employed to classify the solid-state conformation in Csp3 eight-membered rings using data retrieved from an updated version of the Cambridge Structural Database (CSD).

Algorithms↗

MISAE: a new approach for regulatory motif extraction.

The recognition of regulatory motifs of co-regulated genes is essential for understanding the regulatory mechanisms. However, the automatic extraction of regulatory motifs from a given data set of the upstream non-coding DNA sequences of a family of co-regulated genes is difficult because regulatory motifs are often subtle and inexact. This problem is further complicated by the corruption of the data sets. In this paper, a new approach called Mismatch-allowed Probabilistic Suffix Tree Motif Extraction (MISAE) is proposed. It combines the mismatch-allowed probabilistic suffix tree that is a probabilistic model and local prediction for the extraction of regulatory motifs. The proposed approach is tested on 15 co-regulated gene families and compares favorably with other state-of-the-art approaches. Moreover, MISAE performs well on "corrupted" data sets. It is able to extract the motif from a "corrupted" data set with less than one fourth of the sequences containing the real motif.

Algorithms↗

Essential latent knowledge for protein-protein interactions: analysis by an unsupervised learning approach.

Protein-protein interactions play a number of central roles in many cellular functions, including DNA replication, transcription and translation, signal transduction, and metabolic pathways. A recent increase in the number of protein-protein interactions has made predicting unknown protein-protein interactions important for the understanding of living cells. However, the protein-protein interactions experimentally obtained so far are often incomplete and contradictory and, consequently, existing computational prediction methods have integrated evidence (latent knowledge of proteins) from different and more reliable sources. Analyzing the relationships between proteins and the latent knowledge is important to understanding the cellular processes. For this analysis, we propose a new probabilistic model for protein-protein interactions by considering the latent knowledge of proteins. We further present an efficient learning algorithm for this model, based on an EM algorithm. Experimental results have shown that in a supervised test setting, the proposed method outperformed five other competing methods by a statistically significant factor in all cases. Using the probability parameters of a trained model, we have further shown the latent knowledge that is essential to predicting protein-protein interactions. Overall, our experimental results confirm that our proposed model is especially effective for analyzing protein-protein interactions from a viewpoint of the latent knowledge of proteins.

Algorithms↗

Combining sequence and time series expression data to learn transcriptional modules.

Our goal is to cluster genes into transcriptional modules--sets of genes where similarity in expression is explained by common regulatory mechanisms at the transcriptional level. We want to learn modules from both time series gene expression data and genome-wide motif data that are now readily available for organisms such as S. cereviseae as a result of prior computational studies or experimental results. We present a generative probabilistic model for combining regulatory sequence and time series expression data to cluster genes into coherent transcriptional modules. Starting with a set of motifs representing known or putative regulatory elements (transcription factor binding sites) and the counts of occurrences of these motifs in each gene's promoter region, together with a time series expression profile for each gene, the learning algorithm uses expectation maximization to learn module assignments based on both types of data. We also present a technique based on the Jensen-Shannon entropy contributions of motifs in the learned model for associating the most significant motifs to each module. Thus, the algorithm gives a global approach for associating sets of regulatory elements to "modules" of genes with similar time series expression profiles. The model for expression data exploits our prior belief of smooth dependence on time by using statistical splines and is suitable for typical time course data sets with relatively few experiments. Moreover, the model is sufficiently interpretable that we can understand how both sequence data and expression data contribute to the cluster assignments, and how to interpolate between the two data sources. We present experimental results on the yeast cell cycle to validate our method and find that our combined expression and motif clustering algorithm discovers modules with both coherent expression and similar motif patterns, including binding motifs associated to known cell cycle transcription factors.

Algorithms↗

Probabilistic independent component analysis for functional magnetic resonance imaging.

We present an integrated approach to probabilistic independent component analysis (ICA) for functional MRI (FMRI) data that allows for nonsquare mixing in the presence of Gaussian noise. In order to avoid overfitting, we employ objective estimation of the amount of Gaussian noise through Bayesian analysis of the true dimensionality of the data, i.e., the number of activation and non-Gaussian noise sources. This enables us to carry out probabilistic modeling and achieves an asymptotically unique decomposition of the data. It reduces problems of interpretation, as each final independent component is now much more likely to be due to only one physical or physiological process. We also describe other improvements to standard ICA, such as temporal prewhitening and variance normalization of timeseries, the latter being particularly useful in the context of dimensionality reduction when weak activation is present. We discuss the use of prior information about the spatiotemporal nature of the source processes, and an alternative-hypothesis testing approach for inference, using Gaussian mixture models. The performance of our approach is illustrated and evaluated on real and artificial FMRI data, and compared to the spatio-temporal accuracy of results obtained from classical ICA and GLM analyses.

Algorithms↗

Discriminative components of data.

A simple probabilistic model is introduced to generalize classical linear discriminant analysis (LDA) in finding components that are informative of or relevant for data classes. The components maximize the predictability of the class distribution which is asymptotically equivalent to 1) maximizing mutual information with the classes, and 2) finding principal components in the so-called learning or Fisher metrics. The Fisher metric measures only distances that are relevant to the classes, that is, distances that cause changes in the class distribution. The components have applications in data exploration, visualization, and dimensionality reduction. In empirical experiments, the method outperformed, in addition to more classical methods, a Renyi entropy-based alternative while having essentially equivalent computational cost.

Algorithms↗

MCMC data association and sparse factorization updating for real time multitarget tracking with merged and multiple measurements.

In several multitarget tracking applications, a target may return more than one measurement per target and interacting targets may return multiple merged measurements between targets. Existing algorithms for tracking and data association, initially applied to radar tracking, do not adequately address these types of measurements. Here, we introduce a probabilistic model for interacting targets that addresses both types of measurements simultaneously. We provide an algorithm for approximate inference in this model using a Markov chain Monte Carlo (MCMC)-based auxiliary variable particle filter. We Rao-Blackwellize the Markov chain to eliminate sampling over the continuous state space of the targets. A major contribution of this work is the use of sparse least squares updating and downdating techniques, which significantly reduce the computational cost per iteration of the Markov chain. Also, when combined with a simple heuristic, they enable the algorithm to correctly focus computation on interacting targets. We include experimental results on a challenging simulation sequence. We test the accuracy of the algorithm using two sensor modalities, video, and laser range data. We also show the algorithm exhibits real time performance on a conventional PC.

Algorithms↗

Context-based segmentation of image sequences.

We describe an algorithm for context-based segmentation of visual data. New frames in an image sequence (video) are segmented based on the prior segmentation of earlier frames in the sequence. The segmentation is performed by adapting a probabilistic model learned on previous frames, according to the content of the new frame. We utilize the maximum a posteriori version of the EM algorithm to segment the new image. The Gaussian mixture distribution that is used to model the current frame is transformed into a conjugate-prior distribution for the parametric model describing the segmentation of the new frame. This semisupervised method improves the segmentation quality and consistency and enables a propagation of segments along the segmented images. The performance of the proposed approach is illustrated on both simulated and real image data.

Algorithms↗

Probabilistic analysis of the inadvertent reentry of the Cassini spacecraft's radioisotope thermoelectric generators.

As part of the launch approval process, the Interagency Nuclear Safety Review Panel provides an independent safety assessment of space missions--such as the Cassini mission--that carry a significant amount of nuclear materials. This survey article describes potential accident scenarios that might lead to release of fuel from an accidental reentry during an Earth swingby maneuver, the probabilities of such scenarios, and their consequences. To illustrate the nature of calculations used in this area, examples are presented of probabilistic models to obtain both the probability of scenario events and the resultant source terms of such scenarios. Because of large extrapolations from the current knowledge base, the analysis emphasizes treatment of uncertainties.

Algorithms↗

Optimal predictions in everyday cognition.

Human perception and memory are often explained as optimal statistical inferences that are informed by accurate prior probabilities. In contrast, cognitive judgments are usually viewed as following error-prone heuristics that are insensitive to priors. We examined the optimality of human cognition in a more realistic context than typical laboratory studies, asking people to make predictions about the duration or extent of everyday phenomena such as human life spans and the box-office take of movies. Our results suggest that everyday cognitive judgments follow the same optimal statistical principles as perception and memory, and reveal a close correspondence between people's implicit probabilistic models and the statistics of the world.

Cognition↗

Exposure assessment: then, now, and quantum leaps in the future.

Health risk assessments have become so widely accepted in the United States that their conclusions are a major factor in many environmental decisions. Although the risk assessment paradigm is 10 years old, the basic risk assessment process has been used by certain regulatory agencies for nearly 40 years. Each of the four components of the paradigm has undergone significant refinements, particularly during the last 5 years. A recent step in the development of the exposure assessment component can be found in the 1992 EPA Guidelines for Exposure Assessment. Rather than assuming worst-case or hypothetical maximum exposures, these guidelines are designed to lead to an accurate characterization, making use of a number of scientific advances. Many exposure parameters have become better defined, and more sensitive techniques now exist for measuring concentrations of contaminants in the environment. Statistical procedures for characterizing variability, using Monte Carlo or similar approaches, eliminate the need to select point estimates for all individual exposure parameters. These probabilistic models can more accurately characterize the full range of exposures that may potentially be encountered by a given population at a particular site, reducing the need to select highly conservative values to account for this form of uncertainty in the exposure estimate. Lastly, our awareness of the uncertainties in the exposure assessment as well as our knowledge as to how best to characterize them will almost certainly provide evaluations that will be more credible and, therein, more useful to risk managers.(ABSTRACT TRUNCATED AT 250 WORDS)

Environmental Exposure↗