Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

RNA sequence analysis using covariance models.

We describe a general approach to several RNA sequence analysis problems using probabilistic models that flexibly describe the secondary structure and primary sequence consensus of an RNA sequence family. We call these models 'covariance models'. A covariance model of tRNA sequences is an extremely sensitive and discriminative tool for searching for additional tRNAs and tRNA-related sequences in sequence databases. A model can be built automatically from an existing sequence alignment. We also describe an algorithm for learning a model and hence a consensus secondary structure from initially unaligned example sequences and no prior structural information. Models trained on unaligned tRNA examples correctly predict tRNA secondary structure and produce high-quality multiple alignments. The approach may be applied to any family of small RNA sequences.

Algorithms↗

A model for intra-familial distribution of an infectious disease (Chagas' disease).

A probabilistic model for intra-familial distribution of infectious disease is proposed and applied to the prevalence of positive serology for Trypanosoma cruzi infection in a Northeastern Brazilian sample. This double binomial with one tail excess model fits satisfactorily to the data and its interpretation says that around 51% of these 982 families are free of infection risk; among those that are at risk, 3% have a high risk (0.66), probably due to high domestic infestation of the vector bug; while 97% show a small risk (0.11), probably due to accidental, non-domestic transmission.

Binomial Distribution↗

Algorithms for optical mapping.

Optical mapping is a novel technique for determining the restriction sites on a DNA molecule by directly observing a number of partially digested copies of the molecule under a light microscope. The problem is complicated by uncertainty as to the orientation of the molecules and by erroneous detection of cuts. In this paper we study the problem of constructing a restriction map based on optical mapping data. We give several variants of a polynomial reconstruction algorithm, as well as an algorithm that is exponential in the number of cut sites, and hence is appropriate only for small number of cut sites. We give a simple probabilistic model for data generation and for the errors and prove probabilistic upper and lower bounds on the number of molecules needed by each algorithm in order to obtain a correct map, expressed as a function of the number of cut sites and the error parameters. To the best of our knowledge, this is the first probabilistic analysis of algorithms for the problem. We also provide experimental results confirming that our algorithms are highly effective on simulated data.

Algorithms↗

Softening constraints in constraint-based protein topology prediction.

This paper is concerned with the handling of uncertain data about the applicability of constraints in protein topology prediction. It discusses the use of novel methods of representing and reasoning with uncertain data, and presents the results of some experiments in using these methods to build probabilistic models of constraint application. It thus builds on work by other authors in both constraint satisfaction and probabilistic reasoning.

Models, Molecular↗

Assessing the natural attenuation of organic contaminants in aquifers using plume-scale electron and carbon balances: model development with analysis of uncertainty and parameter sensitivity.

A quantitative methodology is described for the field-scale performance assessment of natural attenuation using plume-scale electron and carbon balances. This provides a practical framework for the calculation of global mass balances for contaminant plumes, using mass inputs from the plume source, background groundwater and plume residuals in a simplified box model. Biodegradation processes and reactions included in the analysis are identified from electron acceptors, electron donors and degradation products present in these inputs. Parameter values used in the model are obtained from data acquired during typical site investigation and groundwater monitoring studies for natural attenuation schemes. The approach is evaluated for a UK Permo-Triassic Sandstone aquifer contaminated with a plume of phenolic compounds. Uncertainty in the model predictions and sensitivity to parameter values was assessed by probabilistic modelling using Monte Carlo methods. Sensitivity analyses were compared for different input parameter probability distributions and a base case using fixed parameter values, using an identical conceptual model and data set. Results show that consumption of oxidants by biodegradation is approximately balanced by the production of CH4 and total dissolved inorganic carbon (TDIC) which is conserved in the plume. Under this condition, either the plume electron or carbon balance can be used to determine contaminant mass loss, which is equivalent to only 4% of the estimated source term. This corresponds to a first order, plume-averaged, half-life of > 800 years. The electron balance is particularly sensitive to uncertainty in the source term and dispersive inputs. Reliable historical information on contaminant spillages and detailed site investigation are necessary to accurately characterise the source term. The dispersive influx is sensitive to variability in the plume mixing zone width. Consumption of aqueous oxidants greatly exceeds that of mineral oxidants in the plume, but electron acceptor supply is insufficient to meet the electron donor demand and the plume will grow. The aquifer potential for degradation of these contaminants is limited by high contaminant concentrations and the supply of bioavailable electron acceptors. Natural attenuation will increase only after increased transport and dilution.

Biodegradation, Environmental↗

Stochastic modeling of RNA pseudoknotted structures: a grammatical approach.

MOTIVATION: Modeling RNA pseudoknotted structures remains challenging. Methods have previously been developed to model RNA stem-loops successfully using stochastic context-free grammars (SCFG) adapted from computational linguistics; however, the additional complexity of pseudoknots has made modeling them more difficult. Formally a context-sensitive grammar is required, which would impose a large increase in complexity. RESULTS: We introduce a new grammar modeling approach for RNA pseudoknotted structures based on parallel communicating grammar systems (PCGS). Our new approach can specify pseudoknotted structures, while avoiding context-sensitive rules, using a single CFG synchronized with a number of regular grammars. Technically, the stochastic version of the grammar model can be as simple as an SCFG. As with SCFG, the new approach permits automatic generation of a single-RNA structure prediction algorithm for each specified pseudoknotted structure model. This approach also makes it possible to develop full probabilistic models of pseudoknotted structures to allow the prediction of consensus structures by comparative analysis and structural homology recognition in database searches.

Algorithms↗

A model for the evolution of paralog families in genomes.

We introduce and analyse a simple probabilistic model of genome evolution. It is based on three fundamental evolutionary events: gene loss, duplication and accumulated change. This is motivated by previous works which consisted in fitting the available genomic data into, what is called paralog distributions. This formalism is described by a system of infinite number of linear equations. We show that this system generates a semigroup of linear operators on the space l (1). We prove that size distribution of paralogous gene families in a genome converges to the equilibrium as time goes to infinity. Moreover we show that when probabilities of gene removal and duplication are close to each other, then the resulting distribution is close to logarithmic distribution. Some empirical results for yeast genomes are presented.

Evolution, Molecular↗

Modeling of combined processing steps for reducing Escherichia coli O157:H7 populations in apple cider.

Probabilistic models were used as a systematic approach to describe the response of Escherichia coli O157:H7 populations to combinations of commonly used preservation methods in unpasteurized apple cider. Using a complete factorial experimental design, the effect of pH (3. 1 to 4.3), storage temperature and time (5 to 35 degrees C for 0 to 6 h or 12 h), preservatives (0, 0.05, or 0.1% potassium sorbate or sodium benzoate), and freeze-thaw (F-T; -20 degrees C, 48 h and 4 degrees C, 4 h) treatment combinations (a total of 1,600 treatments) on the probability of achieving a 5-log(10)-unit reduction in a three-strain E. coli O157:H7 mixture in cider was determined. Using logistic regression techniques, pH, temperature, time, and concentration were modeled in separate segments of the data set, resulting in prediction equations for: (i) no preservatives, before F-T; (ii) no preservatives, after F-T; (iii) sorbate, before F-T; (iv) sorbate, after F-T; (v) benzoate, before F-T; and (vi) benzoate, after F-T. Statistical analysis revealed a highly significant (P < 0.0001) effect of all four variables, with cider pH being the most important, followed by temperature and time, and finally by preservative concentration. All models predicted 92 to 99% of the responses correctly. To ensure safety, use of the models is most appropriate at a 0.9 probability level, where the percentage of false positives, i.e., falsely predicting a 5-log(10)-unit reduction, is the lowest (0 to 4.4%). The present study demonstrates the applicability of logistic regression approaches to describing the effectiveness of multiple treatment combinations in pathogen control in cider making. The resulting models can serve as valuable tools in designing safe apple cider processes.

Beverages↗

Career risk of hepatitis C virus infection among U.S. emergency medical and public safety workers.

OBJECTIVE: A probabilistic model was used to analyze the cumulative risk of occupational hepatitis C virus (HCV) infection among U.S. public safety workers. METHODS: A model for the career risk of HCV was developed using the frequency of parenteral exposures to blood, the population seroprevalence of HCV, and the risk of seroconversion after exposure. Estimates of key input variables were obtained from published studies. RESULTS: Calculated estimates of the 30-year risk of infection ranged from <0.1% for police, firefighters, and corrections officers to 1.9% among paramedics and emergency department personnel in high-risk communities. Infrequent exposure to high-risk blood seems to present a greater risk of infection than more frequent contact to low-risk populations. CONCLUSIONS: Use of a probabilistic risk assessment model using published data can assist in policy decisions designed to protect the health and safety of workers. Further efforts to document the frequency of occupationally acquired HCV are needed.

Emergency Medical Technicians↗

Independent factor analysis.

We introduce the independent factor analysis (IFA) method for recovering independent hidden sources from their observed mixtures. IFA generalizes and unifies ordinary factor analysis (FA), principal component analysis (PCA), and independent component analysis (ICA), and can handle not only square noiseless mixing but also the general case where the number of mixtures differs from the number of sources and the data are noisy. IFA is a two-step procedure. In the first step, the source densities, mixing matrix, and noise covariance are estimated from the observed data by maximum likelihood. For this purpose we present an expectation-maximization (EM) algorithm, which performs unsupervised learning of an associated probabilistic model of the mixing situation. Each source in our model is described by a mixture of gaussians; thus, all the probabilistic calculations can be performed analytically. In the second step, the sources are reconstructed from the observed data by an optimal nonlinear estimator. A variational approximation of this algorithm is derived for cases with a large number of sources, where the exact algorithm becomes intractable. Our IFA algorithm reduces to the one for ordinary FA when the sources become gaussian, and to an EM algorithm for PCA in the zero-noise limit. We derive an additional EM algorithm specifically for noiseless IFA. This algorithm is shown to be superior to ICA since it can learn arbitrary source densities from the data. Beyond blind separation, IFA can be used for modeling multidimensional data by a highly constrained mixture of gaussians and as a tool for nonlinear signal encoding.

Algorithms↗

Evaluation of several lightweight stochastic context-free grammars for RNA secondary structure prediction.

BACKGROUND: RNA secondary structure prediction methods based on probabilistic modeling can be developed using stochastic context-free grammars (SCFGs). Such methods can readily combine different sources of information that can be expressed probabilistically, such as an evolutionary model of comparative RNA sequence analysis and a biophysical model of structure plausibility. However, the number of free parameters in an integrated model for consensus RNA structure prediction can become untenable if the underlying SCFG design is too complex. Thus a key question is, what small, simple SCFG designs perform best for RNA secondary structure prediction? RESULTS: Nine different small SCFGs were implemented to explore the tradeoffs between model complexity and prediction accuracy. Each model was tested for single sequence structure prediction accuracy on a benchmark set of RNA secondary structures. CONCLUSIONS: Four SCFG designs had prediction accuracies near the performance of current energy minimization programs. One of these designs, introduced by Knudsen and Hein in their PFOLD algorithm, has only 21 free parameters and is significantly simpler than the others.

Computational Biology↗

Making sense of sparse rating data in collaborative filtering via topographic organization of user preference patterns.

We introduce topographic versions of two latent class models (LCM) for collaborative filtering. Latent classes are topologically organized on a square grid. Topographic organization of latent classes makes orientation in rating/preference patterns captured by the latent classes easier and more systematic. The variation in film rating patterns is modelled by multinomial and binomial distributions with varying independence assumptions. In the first stage of topographic LCM construction, self-organizing maps with neural field organized according to the LCM topology are employed. We apply our system to a large collection of user ratings for films. The system can provide useful visualization plots unveiling user preference patterns buried in the data, without loosing potential to be a good recommender model. It appears that multinomial distribution is most adequate if the model is regularized by tight grid topologies. Since we deal with probabilistic models of the data, we can readily use tools from probability and information theories to interpret and visualize information extracted by our system.

Behavior↗

A critique of Oaksford, Chater, and Larkin's (2000) conditional probability model of conditional reasoning.

M. Oaksford, N. Chater, and J. Larkin (2000) proffered a Bayesian model in which conditional inferences are a direct function of conditional probabilities. In the current article, the authors first considered this model regarding the processing of negatives in conditional reasoning. Its predictions were evaluated against a large-scale meta-analysis (W. J. Schroyens, W. Schaeken, & G. d'Ydewalle, 2001b). This evaluation shows that the model is flawed: The relative size of the negative effects does not match predictions. Next, the authors evaluated the model in relation to inferences about affirmative conditionals, again considering the results of a meta-analysis (W. J. Schroyens, W. Schaeken, & G. d'Ydewalle, 2001a). The conditional probability model is countered by the data reported in literature; a mental models based model produces a better fit. The authors conclude that a purely probabilistic model is deficient and incomplete and cannot do without algorithmic processing assumptions if it is to advance toward a descriptively adequate psychological theory.

Conditioning, Psychological↗

Modelling haemophilia epidemiology and treatment modalities to estimate the unconstrained factor VIII demand.

The article presents a new method for estimating the unconstrained factor VIII (FVIII) demand based on the principles of decision analysis. Epidemiology and treatment modalities were integrated into a model for unconstrained FVIII demand. Assumptions for each variable with impact on the unconstrained FVIII demand were defined and probability estimates for these variables were obtained from the literature and medical experts. The sensitivity of the unconstrained FVIII demand to each of the variables was determined, and the variables with the greatest impact were modelled probabilistically. The probability-weighted average for the unconstrained FVIII demand model was 6.9 units per capita with a 90% uncertainty interval of 2.7-13.6 units per capita. When compared with FVIII usage in countries, only Luxembourg's use of FVIII (7.7 units per capita) exceeded the probability-weighted average for the modelled unconstrained FVIII demand. As better information becomes available, revision of model variables is easily accomplished allowing for a more accurate and dynamic forecast of demand over time. More accurate modelling of the 'true' demand longitudinally should help prevent shortages of FVIII concentrates such as those that have occurred in the past. In addition, a more accurate forecast of FVIII demand will allow national health care policy makers to better allocate financial and other resources. Sufficient and consistent supply of FVIII concentrates and appropriate financing of haemophilia care will allow the clinical benefits of more aggressive treatment regimens such as prophylaxis to be realized.

Adolescent↗

The fusion of large scale classified side-scan sonar image mosaics.

This paper presents a unified framework for the creation of classified maps of the seafloor from sonar imagery. Significant challenges in photometric correction, classification, navigation and registration, and image fusion are addressed. The techniques described are directly applicable to a range of remote sensing problems. Recent advances in side-scan data correction are incorporated to compensate for the sonar beam pattern and motion of the acquisition platform. The corrected images are segmented using pixel-based textural features and standard classifiers. In parallel, the navigation of the sonar device is processed using Kalman filtering techniques. A simultaneous localization and mapping framework is adopted to improve the navigation accuracy and produce georeferenced mosaics of the segmented side-scan data. These are fused within a Markovian framework and two fusion models are presented. The first uses a voting scheme regularized by an isotropic Markov random field and is applicable when the reliability of each information source is unknown. The Markov model is also used to inpaint regions where no final classification decision can be reached using pixel level fusion. The second model formally introduces the reliability of each information source into a probabilistic model. Evaluation of the two models using both synthetic images and real data from a large scale survey shows significant quantitative and qualitative improvement using the fusion approach.

Acoustics↗

A cellular dynamics model of experimental bladder cancer: analysis of the effect of sodium saccharin in the rat.

To make the methodology of risk assessment more consistent with the realities of biological processes, a computer-based model of the carcinogenic process may be used. A previously developed probabilistic model, which is based on a two-stage theory of carcinogenesis, represents urinary bladder carcinogenesis at the cellular level with emphasis on quantification of cell dynamics: cell mitotic rates, cell loss and birth rates, and irreversible cellular transitions from normal to initiated to transformed states are explicitly accounted for. Analyses demonstrate the sensitivity of tumor incidence to the timing and magnitude of changes to these cellular variables. It is demonstrated that response in rats following administration of nongenotoxic compounds, such as sodium saccharin, can be explained entirely on the basis of cytotoxicity and consequent hyperplasia alone.

Animals↗

Evolutionary model selection with a genetic algorithm: a case study using stem RNA.

The choice of a probabilistic model to describe sequence evolution can and should be justified. Underfitting the data through the use of overly simplistic models may miss out on interesting phenomena and lead to incorrect inferences. Overfitting the data with models that are too complex may ascribe biological meaning to statistical artifacts and result in falsely significant findings. We describe a likelihood-based approach for evolutionary model selection. The procedure employs a genetic algorithm (GA) to quickly explore a combinatorially large set of all possible time-reversible Markov models with a fixed number of substitution rates. When applied to stem RNA data subject to well-understood evolutionary forces, the models found by the GA 1) capture the expected overall rate patterns a priori; 2) fit the data better than the best available models based on a priori assumptions, suggesting subtle substitution patterns not previously recognized; 3) cannot be rejected in favor of the general reversible model, implying that the evolution of stem RNA sequences can be explained well with only a few substitution rate parameters; and 4) perform well on simulated data, both in terms of goodness of fit and the ability to estimate evolutionary rates. We also investigate the utility of several distance measures for comparing and contrasting inferred evolutionary models. Using widely available small computer clusters, our approach allows, for the first time, to evaluate the performance of existing RNA evolutionary models by comparing them with a large pool of candidate models and to validate common modeling assumptions. In addition, the new method provides the foundation for rigorous selection and comparison of substitution models for other types of sequence data.

Algorithms↗

Learning higher-order structures in natural images.

The theoretical principles that underlie the representation and computation of higher-order structure in natural images are poorly understood. Recently, there has been considerable interest in using information theoretic techniques, such as independent component analysis, to derive representations for natural images that are optimal in the sense of coding efficiency. Although these approaches have been successful in explaining properties of neural representations in the early visual pathway and visual cortex, because they are based on a linear model, the types of image structure that can be represented are very limited. Here, we present a hierarchical probabilistic model for learning higher-order statistical regularities in natural images. This non-linear model learns an efficient code that describes variations in the underlying probabilistic density. When applied to natural images the algorithm yields coarse-coded, sparse-distributed representations of abstract image properties such as object location, scale and texture. This model offers a novel description of higher-order image structure and could provide theoretical insight into the response properties and computational functions of lower level cortical visual areas.

Learning↗