Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Robust remote homology detection by feature based Profile Hidden Markov Models.

The detection of remote homologies is of major importance for molecular biology applications like drug discovery. The problem is still very challenging even for state-of-the-art probabilistic models of protein families, namely Profile HMMs. In order to improve remote homology detection we propose feature based semi-continuous Profile HMMs. Based on a richer sequence representation consisting of features which capture the biochemical properties of residues in their local context, family specific semi-continuous models are estimated completely data-driven. Additionally, for substantially reducing the number of false predictions an explicit rejection model is estimated. Both the family specific semi-continuous Profile HMM and the non-target model are competitively evaluated. In the experimental evaluation of superfamily based screening of the SCOP database we demonstrate that semi-continuous Profile HMMs significantly outperform their discrete counterparts. Using the rejection model the number of false positive predictions could be reduced substantially which is an important prerequisite for target identification applications.

Journal Article↗

Development of a risk assessment based technique for design/retrofitting of WWTPs.

Up to now, within the design/retrofit of wastewater treatment plants (WWTPs), deterministic models were used to evaluate different scenarios on their merits in terms of effluent compliance. This paper describes an approach in which a Monte Carlo engine is coupled to a deterministic treatment plant model, followed by risk interpretation in the form of concentration-duration-frequency (cdf) curves of norm exceedance. The combination of probabilistic modelling techniques with the currently available deterministic models allows to determine the probability of exceeding the effluent limits of a WWTP. This percentage of exceedance is accompanied by confidence intervals resulting from the inherent uncertainty of influent characteristics and model parameters. The approach is illustrated for a hypothetical case study, consisting of a denitrifying plant model inspired by the benchmark model described by Spanjers et al.

Confidence Intervals↗

A longitudinal model for non-monotonic clinical assessment scale data.

Clinical assessment scales, where subitem ratings are added and summarized as a total score, are convenient tools for monitoring disease progression and often used to measure the effect of drug treatment in clinical trials. Statistical evaluation of any beneficial treatment effects tends to focus on single-valued summary measures, for example, the difference between the score at the end of treatment and the score at baseline. Such analyses ignore potentially important features of the data, e.g. early vs. late recoveries. It is therefore of interest to develop longitudinal models that make more efficient use of the information present in non-monotonic clinical assessment scale data. We propose a two-part modeling approach for the modeling of this type of data. Non-monotonicity is managed by regarding score changes as Markovian transition events. A set of probabilistic models are used to describe the occurrences of the transitions. Continuous models are used to describe the magnitude of the scale score change, given the observed transition. In this manner, a non-monotonic disease progression is handled more efficiently than if other available methods are used. We illustrate this approach using data from a recent phase II study of a drug used in the treatment of stroke, where stroke severity was measured on the Scandinavian Stroke Scale (SSS). This scale consists of nine subitems: consciousness, eye movements, hand/arm/leg motor performance, orientation, speech, facial palsy, and gait. The data were non-monotonic, since there was at any time a risk of a score decline, despite a general tendency towards healing. The two-part probabilistic/continuous model fit the data well and proved to be robust in model-checking procedures such as posterior predictive checks and bootstrapping. The models derived using this approach could potentially accommodate drug effects, not only in terms of score improvement at end of study, but also on the onset of recovery, on dropout and on the probability of unfavorable progression patterns. In addition, it is possible to use the resulting for simulation of the prospective outcome of future studies. We conclude that this approach has considerable potential for more efficient use of information in longitudinal modeling of non-monotonic clinical assessment scale data.

Algorithms↗

Kinetic and dynamic models of diving gases in decompression sickness prevention.

Decompression sickness is a complex phenomenon involving gas exchange, bubble dynamics and tissue response. Relatively simple deterministic compartmental models using empirically derived parameters have been the mainstay of the practice for preventing decompression sickness since the early 1900s. Decades of research have improved our understanding of decompression physiology, and the insights incorporated in decompression models have allowed people to dive deeper into the ocean. However, these efforts have not yet, and are unlikely in the near future, to result in a 'universal' deterministic model that can predict when decompression sickness will occur. Divers using current recreational dive computers need to be aware of their limitations. Probabilistic models based on the estimation of parameters using modern statistical methods from large databases of dives offer a new approach and can provide a means of standardisation of deterministic models. Future improvements in decompression practice will depend on continued improvement in understanding the kinetics and dynamics of gas exchange, bubble evolution and tissue response, and the incorporation of this knowledge in risk models whose parameters can be estimated from large databases of human and animal data.

Decompression Sickness↗

A modeling framework for estimating children's residential exposure and dose to chlorpyrifos via dermal residue contact and nondietary ingestion.

To help address the Food Quality Protection Act of 1996, a physically based probabilistic model has been developed to quantify and analyze dermal and nondietary ingestion exposure and dose to pesticides. The Residential Stochastic Human Exposure and Dose Simulation Model for Pesticides (Residential-SHEDS) simulates the exposures and doses of children contacting residues on surfaces in treated residences and on turf in treated residential yards. The simulations combine sequential time-location-activity information from children's diaries with microlevel videotaped activity data, probability distributions of measured surface residues and exposure factors, and pharmacokinetic rate constants. Model outputs include individual profiles and population statistics for daily dermal loading, mass in the blood compartment, ingested residue via nondietary objects, and mass of eliminated metabolite, as well as contributions from various routes, pathways, and media. To illustrate the capabilities of the model framework, we applied Residential-SHEDS to estimate children's residential exposure and dose to chlorpyrifos for 12 exposure scenarios: 2 age groups (0-4 and 5-9 years); 2 indoor pesticide application methods (broadcast and crack and crevice); and 3 postindoor application time periods (< 1, 1-7, and 8-30 days). Independent residential turf applications (liquid or granular) were included in each of these scenarios. Despite the current data limitations and model assumptions, the case study predicts exposure and dose estimates that compare well to measurements in the published literature, and provides insights to the relative importance of exposure scenarios and pathways.

Administration, Cutaneous↗

[The determination of the sensitivity and specificity of the interpretation rules parameters in a study of the medical equipment testing].

The interpretation systems of medical instruments are traditionally assessed by their sensitivity and specificity in definitions of the binary classification scheme. However, should the system's conclusions on the presence of a pathology be extended by a multitude of degrees of its pronouncement, such scheme is unacceptable. A three-digital classification scheme with isolated degrees of pronounced and weak pathology signs reflects, firstly, an understanding of the interpretation error value between these two degrees and, secondly, the harmlessness of errors inside each degree. The discussed logical and probabilistic models of constructing the quality parameters ensure the continuity in respect to the indices of the binary classification scheme, however, it can be stated that the probabilistic parameters are more preferable.

Data Interpretation, Statistical↗

Statistical mechanics of the Bayesian image restoration under spatially correlated noise.

We investigated the use of the Bayesian inference to restore noise-degraded images under conditions of spatially correlated noise. The generative statistical models used for the original image and the noise were assumed to obey multidimensional Gaussian distributions, whose covariance matrices are translational invariant. We derived an exact description to be used as the expectation for the restored image by the Fourier transformation and restored an image distorted by spatially correlated noise by using a spatially uncorrelated noise model. We found that the resulting hyperparameter estimations for the minimum error and maximal posterior marginal criteria did not coincide when the generative probabilistic model and the model used for restoration were in different classes, while they did coincide when they were in the same class.

Journal Article↗

Eye-movement models for arithmetic and reading performance.

Three stochastic eye-movement models for arithmetic and reading performance have been proposed, one for arithmetic and two for reading. Each model characterizes a real-time stochastic process in terms of fixation durations and saccadic movement, but only direction and length of saccades are considered, not acceleration or velocity. Aspects of the models that are emphasized, partly because of their general neglect in the literature, are the probability distribution of fixation durations and the random walk of saccade directions. The distributions of fixation duration are approximately exponential, but systematic deviations can be accounted for in the models, even though the fit to data is not perfect. In the case of the arithmetic algorithms of addition and subtraction, the random walk of the normative model has only two possible moves. Data are also presented on backtracking, skipping and wandering eye movements, each of which has a significant relative frequency. The first reading model is called a minimal control model, because it does not take account of the effects of many local variables, e.g., word length, that have been extensively studied. The axioms on fixation duration for the minimal control model are the same as for the arithmetic model. Abstracting from the different arrangement of stimuli in arithmetic algorithms and in linear text, the axioms on saccadic motion for the two models are also essentially identical. The stochastic nature of both models is strongly supported by data on the independence of fixation durations from previous fixation durations. Additional detailed evidence is presented for the arithmetic model. To better account for a great variety of experimental results concerning significant effects on eye movements in reading, a text-dependent probabilistic model of reading is introduced. Significant local effects fall into three classes, identified as line, word and grammatical variables. The revised axioms embody five features of text known to be significant: (i) fixation duration depends on the number of letters in a word; (ii) a saccade is longer when a longer word is to the right; (iii) a saccade is longer when the current fixation is on a longer word; (iv) high-frequency fixation words have the highest probability of being skipped; (v) ambiguous or difficult grammatical structures increase backtracking.

Cognition↗

Modelling the determinants of trans-Tasman migration after World War II.

"This paper identifies the economic and demographic factors responsible for migration flows between Australia and New Zealand by means of a probabilistic model of emigration in both directions. The largely uncontrolled flows between the two countries have the same determinants as those commonly found in studies of internal migration. The cost of migration (proxied by the real cost of air travel), labour market conditions and the potential earnings differential play a role, although the results are modified by the incidence of return migration and age composition."

Age Factors↗

Probabilistic evaluation of aspirating doses of a Q fever pathogen.

The process of biological aerosol penetration into respiratory organs is connected with the estimation of the amount of an aspiration dose of microorganisms (D) a human gets in the center of infection. Here we submit a probabilistic model describing the process of hits of a Q fever pathogen in a human respiratory tract. This approach makes it possible to get qualitatively different probabilistic estimations of doses received and corresponding values of critical time intervals of exposure at which the amount accumulated in the respiratory tract turns into an infecting dose. This permits one to approach the explanation of the infection process in the context of the absence of its absolute nature, which is connected both with the immune system responses and the irregularity with which the recipient receives aspiration doses under conditions of artificial distribution of the microbe-bearing aerosol.

Aerosols↗

"Within patient"-dependent outcomes in graft occlusion after coronary artery bypass. SINBA Group.

In clinical trials aimed at assessing the efficacy of drugs after coronary artery grafting, statistical analysis is usually carried out on a per-patient basis and on a per-graft basis. Owing to the independence of responses among patients, the former outcome is analyzed by assuming a binomial model. However, this model cannot be directly adopted in analyzing the latter outcome because multiple vein grafts in the same patient do not act independently. It has been suggested that patients be considered as clusters of distal anastomoses, subsequently comparing the frequencies of occlusion under different treatments after adjusting the variance of each frequency for the size of the cluster. Alternatively, one can measure the intraclass (intrapatient) correlation coefficient and adopt distribution functions that include this quantity as one of the parameters. This approach has been adopted by many authors although the studies differed in the statistical model used to represent this kind of data. Two probability functions, Altham's and beta-binomial, have found wide application in different biomedical fields. Both are able to model the extrabinomial variability since their tails tend to zero more slowly than those of the binomial distribution. An alternative to adopting a specific probabilistic model consists of specifying a mixed model or a Markov-like susceptibility model. These models follow the same rationale since they assume the existence of two processes, one that causes an individual to become susceptible to occlusion, the other that determines the subsequent probability of occlusion in the individual. After presenting these models in a unified framework, this paper compares the estimates obtained by fitting the models to data gathered at coronary angiography 1 year after surgery by the SINBA group on a total of 847 saphenous vein anastomoses in 349 patients. Finally, the issue concerning the contribution of each patient to the effective sample size is discussed.

Coronary Angiography↗

Probabilistic risk assessment for personal exposure to carcinogenic polycyclic aromatic hydrocarbons in Taiwanese temples.

To assess how the human exposure to airborne carcinogenic polycyclic aromatic hydrocarbons (PAHs) during working in or visiting a typical Taiwanese temple, we present a probabilistic risk model, appraised with reported empirical data. Two approaches are applied, one based on animal-derived benzo[a]pyrene (B[a]P) toxic equivalents (B[a]P(eq)) of individual PAHs and one is assumed that the potency of PAH mixtures is linked to their B[a]P level. The model integrates probabilistic exposure profiles of total-PAH and particle-bound PAH levels inside a temple from a published exploratory study with probabilistic incremental lifetime cancer risk (ILCR) models taking into account inhalation and dermal contact pathways, to quantitatively estimate the exposure risks for three age groups of adult, adolescent, and child. Risk analysis indicates that 90% probability inhalation ILCRs for three age groups have orders of magnitude around 10(-7)--10(-6); whereas for the dermal contact ILCRs ranging from 10(-5) to 10(-4), indicating high potential cancer risk. All 90% probabilities of B[a]P- and B[a]P(eq)-based total ILCRs are larger than 10(-6), indicating unacceptable probability distributions for three age groups. Sensitivity analysis indicates that to increase the accuracy of the results efforts should focus on a better definition of probability distributions for inhalation cancer slope factor, inhalation rates, and particle-bound PAH-to-skin adherence factor. We estimate risk-based visiting frequency advice for adult, adolescent, and child to a temple ranging from 5 to 7, 17 to 23, and 48 to 65 year(-1), respectively, based on an average 3h residence time.

Adolescent↗

A latent variable model for chemogenomic profiling.

MOTIVATION: In haploinsufficiency profiling data, pleiotropic genes are often misclassified by clustering algorithms that impose the constraint that a gene or experiment belong to only one cluster. We have developed a general probabilistic model that clusters genes and experiments without requiring that a given gene or drug only appear in one cluster. The model also incorporates the functional annotation of known genes to guide the clustering procedure. RESULTS: We applied our model to the clustering of 79 chemogenomic experiments in yeast. Known pleiotropic genes PDR5 and MAL11 are more accurately represented by the model than by a clustering procedure that requires genes to belong to a single cluster. Drugs such as miconazole and fenpropimorph that have different targets but similar off-target genes are clustered more accurately by the model-based framework. We show that this model is useful for summarizing the relationship among treatments and genes affected by those treatments in a compendium of microarray profiles. AVAILABILITY: Supplementary information and computer code at http://genomics.lbl.gov/llda.

Computer Simulation↗

Global mapping of pharmacological space.

We present the global mapping of pharmacological space by the integration of several vast sources of medicinal chemistry structure-activity relationships (SAR) data. Our comprehensive mapping of pharmacological space enables us to identify confidently the human targets for which chemical tools and drugs have been discovered to date. The integration of SAR data from diverse sources by unique canonical chemical structure, protein sequence and disease indication enables the construction of a ligand-target matrix to explore the global relationships between chemical structure and biological targets. Using the data matrix, we are able to catalog the links between proteins in chemical space as a polypharmacology interaction network. We demonstrate that probabilistic models can be used to predict pharmacology from a large knowledge base. The relationships between proteins, chemical structures and drug-like properties provide a framework for developing a probabilistic approach to drug discovery that can be exploited to increase research productivity.

Computational Biology↗

Analysis of a circular code model.

A circular code has been identified in the protein (coding) genes of both eukaryotes and prokaryotes by using a statistical method called trinucleotide frequency (TF) method [Arquès & Michel (1996). J. theor. Biol. 182, 45-58]. Recently, a probabilistic model based on the nucleotide frequencies with a hypothesis of absence of correlation between successive bases on a DNA strand, has been proposed by Koch & Lehmann [(1997). J. theor. Biol. 189, 171-174] for constructing some particular circular codes. Their interesting method which we call here nucleotide frequency (NF) method, reveals several limits for constructing the circular code observed with protein genes.

Animals↗

A population exposure model for particulate matter: case study results for PM(2.5) in Philadelphia, PA.

A population exposure model for particulate matter (PM), called the Stochastic Human Exposure and Dose Simulation (SHEDS-PM) model, has been developed and applied in a case study of daily PM(2.5) exposures for the population living in Philadelphia, PA. SHEDS-PM is a probabilistic model that estimates the population distribution of total PM exposures by randomly sampling from various input distributions. A mass balance equation is used to calculate indoor PM concentrations for the residential microenvironment from ambient outdoor PM concentrations and physical factor data (e.g., air exchange, penetration, deposition), as well as emission strengths for indoor PM sources (e.g., smoking, cooking). PM concentrations in nonresidential microenvironments are calculated using equations developed from regression analysis of available indoor and outdoor measurement data for vehicles, offices, schools, stores, and restaurants/bars. Additional model inputs include demographic data for the population being modeled and human activity pattern data from EPA's Consolidated Human Activity Database (CHAD). Model outputs include distributions of daily total PM exposures in various microenvironments (indoors, in vehicles, outdoors), and the contribution from PM of ambient origin to daily total PM exposures in these microenvironments. SHEDS-PM has been applied to the population of Philadelphia using spatially and temporally interpolated ambient PM(2.5) measurements from 1992-1993 and 1990 US Census data for each census tract in Philadelphia. The resulting distributions showed substantial variability in daily total PM(2.5) exposures for the population of Philadelphia (median=20 microg/m(3); 90th percentile=59 microg/m(3)). Variability in human activities, and the presence of indoor-residential sources in particular, contributed to the observed variability in total PM(2.5) exposures. The uncertainty in the estimated population distribution for total PM(2.5) exposures was highest at the upper end of the distribution and revealed the importance of including estimates of input uncertainty in population exposure models. The distributions of daily microenvironmental PM(2.5) exposures (exposures due to time spent in various microenvironments) indicated that indoor-residential PM(2.5) exposures (median=13 microg/m(3)) had the greatest influence on total PM(2.5) exposures compared to the other microenvironments. The distribution of daily exposures to PM(2.5) of ambient origin was less variable across the population than the distribution of daily total PM(2.5) exposures (median=7 microg/m(3); 90th percentile=18 microg/m(3)) and similar to the distribution of ambient outdoor PM(2.5) concentrations. This result suggests that human activity patterns did not have as strong an influence on ambient PM(2.5) exposures as was observed for exposure to other PM(2.5) sources. For most of the simulated population, exposure to PM(2.5) of ambient origin contributed a significant percent of the daily total PM(2.5) exposures (median=37.5%), especially for the segment of the population without exposure to environmental tobacco smoke in the residence (median=46.4%). Development of the SHEDS-PM model using the Philadelphia PM(2.5) case study also provided useful insights into the limitations of currently available data for use in population exposure models. In addition, data needs for improving inputs to the SHEDS-PM model, reducing uncertainty and further refinement of the model structure, were identified.

Activities of Daily Living↗

A mammalian promoter model links cis elements to genetic networks.

An accurate identification of gene promoters remains an important challenge. Computational approaches for this problem rely on promoter sequence attributes that are believed to be critical for transcription initiation. Here we report a probabilistic model that captures two important properties of promoters, not used by previous methods, viz., the location preference and co-occurrence of promoter elements. Additionally, we found that many of the position-specific DNA elements are strongly linked with the function of the gene product. For instance, a highly conserved motif CCTTT at -1 position is strongly associated with protein synthesis, cellular and tissue development. Our comparative analysis of promoter classes reveals that the promoters devoid of CpG islands are more conserved and have fewer alternative transcription start sites. The discovered links between promoter elements and gene function allows us to infer genetic networks from promoter elements. The web server for the PSPA promoter predictor is available at /PSPA.

Animals↗

A unified approach to joint modeling of multiple quantitative and qualitative traits in gene mapping.

Using graph theory, we present a theoretical basis for mapping oligogenes in the joint presence of multiple phenotypic measurements of both quantitative and qualitative types. Various statistical models proposed earlier for several traits of solely single type are special cases of the unified approach given here. Our emphasis is on the generality of the framework, without specifying explicit assumptions about a sampling design. When information about environmental factors potentially affecting the traits is available, it can be incorporated into the genetic model. We adopt the Bayesian inferential machinery due to its firm theoretical basis and its capability of handling uncertain quantities; such as unobserved model parameters, missing marker data, and even different putative genetic models, probabilistically within a single framework. It is shown here that biological hypotheses about single gene affecting simultaneously multiple traits (pleiotropy) can be intuitively imposed as parameter constraints, leading to pleiotropic models for which posterior probabilities can be calculated. Outline of the possible implementation of the Bayesian method is described using the general reversible-jump Markov chain Monte Carlo algorithm. Some future challenges and extensions are also discussed.

Animals↗