Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Incorporating exposure models in probabilistic assessment of the risks of premature mortality from particulate matter.

This paper examines the link between the ambient level of particulate pollution and subsequent human health effects and various sources of uncertainty when total exposure is taken into consideration. The exposure simulation model statistically simulates daily personal total exposure to ambient PM and nonambient PM generated from indoor sources. It incorporates outdoor-indoor penetration of PM, contributions of PM from indoor sources, and time-activity patterns for target groups of the population. The model is illustrated for Los Angeles County using recent 1997 monitoring data for both PM(10) and PM(2.5). The results indicate that, on average, outdoor-source PM contributes about 20-25% of the total PM exposure to Los Angeles County individuals not exposed to environmental tobacco smoking (ETS), and about 15% for those who are exposed to ETS. The model computes both the fractional contribution of outdoor concentrations to total exposure and the effect of exposure uncertainties on the estimated slope of the (linear) concentration-response curve in time-series studies for PM health effects. The latter considers the effects of measurement and misclassification error on PM epidemiological time-series studies. The paper compares the predictions of a conventional PM epidemiological model, based solely on ambient concentration measurements at a central monitoring station, and an exposure simulation model, which considers the quantitative relationship between central-monitoring PM concentrations and total individual exposures to particulate matter. The results show that the effects of adjusting from outdoor concentrations to personal exposures and correcting dose-response bias are nearly equal, so that roughly the same premature mortalities associated with short-term exposure to both ambient PM(2.5) and PM(10) in Los Angeles County are predicted with both models. The uncertainty in the slope of the concentration-response curve in the time-series studies is the single most important source of uncertainty in both the ambient- and the exposure-health model.

Air Pollutants↗

Diffusion tensor imaging in primary brain tumors: reproducible quantitative analysis of corpus callosum infiltration and contralateral involvement using a probabilistic mixture model.

Diffusion tensor imaging (DTI) has been advocated as a promising tool for delineation of the extent of tumor infiltration by primary brain tumors. First reports show conflicting results mainly due to difficulties in reproducible determination of DTI-derived parameters. A novel method based on probabilistic voxel classification for a user-independent analysis of DTI-derived parameters is presented and tested in healthy controls and patients with primary brain tumors. The proposed quantification method proved to be highly reproducible both in healthy controls and patients. Fiber integrity in the corpus callosum (CC) was measured using this quantification method, and the profiles of fractional anisotropy (FA) provided additional information of the possible extent of infiltration of primary brain tumors when compared to conventional imaging. This yielded additional information on the nature of ambiguous contralateral lesions in patients with primary brain tumors. The results show that DTI-derived parameters can be determined reproducibly and may have a strong impact on evaluation of contralateral extent of primary brain tumors.

Adult↗

Free-space optical communication through a forest canopy.

We model the effects of the leaves of mature broadleaf (deciduous) trees on air-to-ground free-space optical communication systems operating through the leaf canopy. The concept of leaf area index (LAI) is reviewed and related to a probabilistic model of foliage consisting of obscuring leaves randomly distributed throughout a treetop layer. Individual leaves are opaque. The expected fractional unobscured area statistic is derived as well as the variance around the expected value. Monte Carlo simulation results confirm the predictions of this probabilistic model. To verify the predictions of the statistical model experimentally, a passive optical technique has been used to make measurements of observed sky illumination in a mature broadleaf environment. The results of the measurements, as a function of zenith angle, provide strong evidence for the applicability of the model, and a single parameter fit to the data reinforces a natural connection to LAI. Specific simulations of signal-to-noise ratio degradation as a function of zenith angle in a specific ground-to-unmanned aerial vehicle communication situation have demonstrated the effect of obscuration on performance.

Journal Article↗

[A predictive model for affect of atopic dermatitis in infancy by neural network and multiple logistic regression].

OBJECTS: To analyze the predictive accuracy of the predictive model for affect of atopic dermatitis in infancy, from the data of the epidemiological survey, which were conducted for 10,000 of mothers of infants and children in 1993. SUBJECTS AND METHODS: A total of 4610 replies were received: 2714 from mothers of infants (12 month old) and 1,896 from mothers of children (2 years old). The sensitivity, specificity and predictive accuracy were calculated from probabilistic model by neural network analysis (NNA) and multiple logistic regression analysis (MLA). RESULTS: Risk factors for probabilistic model by NNA were family history (father, mother, siblings, grand father, grand mother), food restriction, food allergy, age, food restriction of mother, egg introduced time, cow's milk introduced time. The sensitivity, specificity and predictive accuracy of NNA model was 88.6%, 99.5% and 96.4%, respectively and MLA model was 75.1%, 82.6% and 82.3%, respectively. CONCLUSION: These results suggest that the NNA is a good and useful method for prediction of onset of AD than MLA. Furthermore, It is necessary to investigate the artificial neural networks for diagnosis and/or treatment by physician.

Child, Preschool↗

Financial consequences of changes in health care demands related to tobacco consumption in Mexico: information for policy makers.

This paper presents the results from a longitudinal study in which the main purpose was to determine the health-care costs and financial consequences of changes in the health care demands related to tobacco consumption in Mexico. Eleven health interventions were selected to conduct this study and four probabilistic models were developed to forecast the expected changes in the epidemiologic profile of selected diseases. The costing method was based on the identification of case management costs using the instrumentation and consensus techniques, probabilistic models were designed using the Box-Jenkins technique and allowed us to identify the expected case trends for the 2001-2003 period. The generation of information on case management costs for the selected interventions is a central instrument in the planning of health programs, above all in that which refers to resource allocation by type of demand. On the other hand, the identification of expected cases and the financial consequences allowed us to know the growing trends of the sums required to satisfy health care demands for the period under study. The three types of information are a relevant resource for decision-makers in the production and financing of health services.

Case Management↗

Bayesian segmentation of protein secondary structure.

We present a novel method for predicting the secondary structure of a protein from its amino acid sequence. Most existing methods predict each position in turn based on a local window of residues, sliding this window along the length of the sequence. In contrast, we develop a probabilistic model of protein sequence/structure relationships in terms of structural segments, and formulate secondary structure prediction as a general Bayesian inference problem. A distinctive feature of our approach is the ability to develop explicit probabilistic models for alpha-helices, beta-strands, and other classes of secondary structure, incorporating experimentally and empirically observed aspects of protein structure such as helical capping signals, side chain correlations, and segment length distributions. Our model is Markovian in the segments, permitting efficient exact calculation of the posterior probability distribution over all possible segmentations of the sequence using dynamic programming. The optimal segmentation is computed and compared to a predictor based on marginal posterior modes, and the latter is shown to provide significant improvement in predictive accuracy. The marginalization procedure provides exact secondary structure probabilities at each sequence position, which are shown to be reliable estimates of prediction uncertainty. We apply this model to a database of 452 nonhomologous structures, achieving accuracies as high as the best currently available methods. We conclude by discussing an extension of this framework to model nonlocal interactions in protein structures, providing a possible direction for future improvements in secondary structure prediction accuracy.

Algorithms↗

Projected number of diabetic renal disease patients among insulin-dependent diabetes mellitus children in Japan using a Markov model with probabilistic sensitivity analysis.

BACKGROUND: To plan prevention programmes for the diabetic renal disease among insulin-dependent diabetes mellitus (IDDM) children, projections of future trends for the disease is crucial. We projected future trends in the number of diabetic renal disease patients among IDDM children and assessed an impact of treatment dissemination in Japan. METHODS: We used a Markov model to describe the clinical courses of diabetic renal disease. Future trends in the number of patients with diabetic nephropathy (DN) and end-stage renal disease (ESRD) were projected from the year 1995 to 2015. We made three scenarios for assessing an impact of the dissemination of new treatment. We performed a probabilistic sensitivity analysis for the uncertainty of transition probabilities. RESULTS: The results showed that the number of patients with DN was 790.5 (5th to 95th percentile: 652.5-955.1), ESRD was 253.3 (5th to 95th percentile: 207.3-310.0) in year 2015 on basic scenario. Considering the dissemination of intensive insulin therapy, under the scenario of the gradual increase of the treatment, the result showed that the number of patients with DN was 713.1 (5th to 95th percentile: 546.2-930.6), ESRD was 231.0 (5th to 95th percentile: 176.6-296.2). Under the scenario of the immediate change of the treatment, the results showed that the number of patients with DN in 2015 was 418.9 (5th percentile; 345.4; 95th percentile; 506.1) and with ESRD was 133.4 (5th percentile; 109.0; 95th percentile; 163.8). CONCLUSIONS: The results of the projection showed a gradual increase in the number of patients with DMN and ESRD. Examination of three possible scenarios showed that the programme of dissemination of intensive insulin therapy prevented the progression of diabetic renal disease.

Adolescent↗

Understanding tuberculosis epidemiology using structured statistical models.

Molecular epidemiological studies can provide novel insights into the transmission of infectious diseases such as tuberculosis. Typically, risk factors for transmission are identified using traditional hypothesis-driven statistical methods such as logistic regression. However, limitations become apparent in these approaches as the scope of these studies expand to include additional epidemiological and bacterial genomic data. Here we examine the use of Bayesian models to analyze tuberculosis epidemiology. We begin by exploring the use of Bayesian networks (BNs) to identify the distribution of tuberculosis patient attributes (including demographic and clinical attributes). Using existing algorithms for constructing BNs from observational data, we learned a BN from data about tuberculosis patients collected in San Francisco from 1991 to 1999. We verified that the resulting probabilistic models did in fact capture known statistical relationships. Next, we examine the use of newly introduced methods for representing and automatically constructing probabilistic models in structured domains. We use statistical relational models (SRMs) to model distributions over relational domains. SRMs are ideally suited to richly structured epidemiological data. We use a data-driven method to construct a statistical relational model directly from data stored in a relational database. The resulting model reveals the relationships between variables in the data and describes their distribution. We applied this procedure to the data on tuberculosis patients in San Francisco from 1991 to 1999, their Mycobacterium tuberculosis strains, and data on contact investigations. The resulting statistical relational model corroborated previously reported findings and revealed several novel associations. These models illustrate the potential for this approach to reveal relationships within richly structured data that may not be apparent using conventional statistical approaches. We show that Bayesian methods, in particular statistical relational models, are an important tool for understanding infectious disease epidemiology.

Adult↗

Maxwell and very-hard-particle models for probabilistic ballistic annihilation: hydrodynamic description.

The hydrodynamic description of probabilistic ballistic annihilation, for which no conservation laws hold, is an intricate problem with hard spherelike dynamics for which no exact solution exists. We consequently focus on simplified approaches, the Maxwell and very-hard-particle (VHP) models, which allows us to compute analytically upper and lower bounds for several quantities. The purpose is to test the possibility of describing such a far from equilibrium dynamics with simplified kinetic models. The motivation is also in turn to assess the relevance of some singular features appearing within the original model and the approximations invoked to study it. The scaling exponents are first obtained from the (simplified) Boltzmann equation, and are confronted against direct Monte Carlo simulations. Then, the Chapman-Enskog method is used to obtain constitutive relations and transport coefficients. The corresponding Navier-Stokes equations for the hydrodynamic fields are derived for both Maxwell and VHP models. We finally perform a linear stability analysis around the homogeneous solution, which illustrates the importance of dissipation in the possible development of spatial inhomogeneities.

Journal Article↗

Assessing patient safety risk before the injury occurs: an introduction to sociotechnical probabilistic risk modelling in health care.

Since 1 July 2001 the Joint Commission on Accreditation of Healthcare Organizations (JCAHO) has required each accredited hospital to conduct at least one proactive risk assessment annually. Failure modes and effects analysis (FMEA) was recommended as one tool for conducting this task. This paper examines the limitations of FMEA and introduces a second tool used by the aviation and nuclear industries to examine low frequency, high impact events in complex systems. The adapted tool, known as sociotechnical probabilistic risk assessment (ST-PRA), provides an alternative for proactively identifying, prioritizing, and mitigating patient safety risk. The uniqueness of ST-PRA is its ability to model combinations of equipment failures, human error, at risk behavioral norms, and recovery opportunities through the use of fault trees. While ST-PRA is a complex, high end risk modelling tool, it provides an opportunity to visualize system risk in a manner that is not possible through FMEA.

Hospital Administration↗

Handling of contamination variability in exposure assessment: a case study with ochratoxin A.

The contamination of foods dedicated to human consumption varies over space and time. In exposure assessment, this is usually addressed through probabilistic modelling. The present work explores how the variability and uncertainty of exposures estimated at the population level are affected by: (a) the (non-)parametric nature of input contamination distributions; (b) the time-window used to sample contamination values within those distributions. Focusing on exposure of the French population to food mycotoxin ochratoxin A, we implement a range of second-order Monte-Carlo simulations that allow distinguishing variability of exposures from uncertainty of distributional parameters estimates. A simulation runs 10,000 iterations. Overall estimates of parameters are given by the median across iterations and 95%CI by 2.5th and 97.5th percentiles. Our results show that: (a) parametric (log-normal) input distributions may lead to over-estimation of variability and greater uncertainty as compared to non-parametric ones (P97.5 [95%CI] of 7.1 [6.6;7.7] for Parametric-Occasion, 4.6 [4.3;5.0] for Non-Parametric-Occasion), and that (b) the 'Occasion' time-window combines better estimate of variability and lower uncertainty when exposure modelling is applied to populations living in developed countries with complex agri-food systems (P97.5 [95%CI]: 7.3 [6.2;8.9] for Non-Parametric-Week, 4.6 [4.3;5.0] for Non-Parametric-Occasion). A deterministic approach is nevertheless preferred to probabilistic modelling every time input data quality is questionable.

Adolescent↗

Probabilities and predictions: modeling the development of scientific problem-solving skills.

The IMMEX (Interactive Multi-Media Exercises) Web-based problem set platform enables the online delivery of complex, multimedia simulations, the rapid collection of student performance data, and has already been used in several genetic simulations. The next step is the use of these data to understand and improve student learning in a formative manner. This article describes the development of probabilistic models of undergraduate student problem solving in molecular genetics that detailed the spectrum of strategies students used when problem solving, and how the strategic approaches evolved with experience. The actions of 776 university sophomore biology majors from three molecular biology lecture courses were recorded and analyzed. Each of six simulations were first grouped by artificial neural network clustering to provide individual performance measures, and then sequences of these performances were probabilistically modeled by hidden Markov modeling to provide measures of progress. The models showed that students with different initial problem-solving abilities choose different strategies. Initial and final strategies varied across different sections of the same course and were not strongly correlated with other achievement measures. In contrast to previous studies, we observed no significant gender differences. We suggest that instructor interventions based on early student performances with these simulations may assist students to recognize effective and efficient problem-solving strategies and enhance learning.

Aptitude↗

Crossing the river stone by stone: approaches for residential risk assessment for consumers.

Consumer products may contain constituents that warrant a risk analysis if they raise toxicological concern. Risk assessments are performed a priori, e.g. for pesticides and biocides, and a posteriori, to diagnose risks of contaminants. An overview is presented of residential exposure assessment and risk characterization. For exposure assessment, predictive models are used to estimate exposure concentrations. The available data on product use are used to quantify the intensity of exposure. Often, both exposure concentration and product use show high variability. Worst case assessments cope with variability and uncertainty in data poor situations by selecting 'worst case' values for exposures and exposure factors. Probabilistic models may be used to quantify and model variability and uncertainty when appropriate data is available. The Margin Of Safety approach to characterize risk is discussed. Many biocides handled by consumers are used now and then and (sub)acute exposure and toxicology will be most relevant. Users and children are generally seen as critical groups during the application and post-application phases of exposure, respectively. Still, the diversity of consumer products requires consideration of the merits of each case. We conclude that residential risk assessment is still searching for methods, data and models. Probabilistic methods appear to be useful tools, but a major challenge is to integrate them in regulatory frameworks.

Adult↗

Approaches to assessment of exposure to food- and supplement-derived amino acids.

Although the amino acid composition of almost all food proteins is known, estimating the amino acid intake from the diet is extremely difficult because of the lack of available data. A conservative approach would be to determine the population distribution of protein intake, select the 97.5(th) or higher percentile of intake, assume all comes from the target protein, and estimate exposure to some specific amino acid. Any number of dietary survey methodologies could be used to conduct such a conservative approach. However, given the great variety of brands of food supplements, estimates of amino acid intakes from supplements are problematic. Firstly, few studies include supplements in their target nutrient sources because brand-level data would need to be retained and nutritional composition data would need to be recorded. Probabilistic modeling offers some solution provided some basic data are gathered. The percentage of the population regularly taking supplements and the frequency of consumption must be known. Therefore, data on the dietary supplement market would need to be known including the percent of brands containing amino acids and if possible specific amino acids together with concentrations. A probabilistic model as follows would ensue: probability of being a consumer of amino acid supplements; probability distribution function of frequency of use of supplements; probability distribution function of dose per eating occasion; market characteristics; probability distribution function for dietary amino acid intake. Using multiple iterations and perhaps bootstrapping on some elements of the model, fully worst-case model scenarios of exposure could be computed.

Amino Acids↗

Cost effectiveness of treatment for amblyopia: an analysis based on a probabilistic Markov model.

AIMS: To estimate the long term cost effectiveness of treatment for amblyopia in 3 year old children. METHODS: A cost utility analysis was performed using decision analysis including a Markov state transition model. Incremental costs and effects during the children's remaining lifetime were estimated. The model took into account the costs and success rate of treatment as well as effects of unilateral and bilateral visual impairment caused by amblyopia and other eye diseases coming along later in life on quality of life (utility). Model parameter values were obtained from the literature, and from a survey of experts. For the utility of unilateral visual impairment a base value of 0.96 was assumed. Costs were estimated from a third party payer perspective for the year 2002 in Germany. Costs and effects were discounted at 3%. Uncertainty was assessed by univariate and probabilistic sensitivity analysis (Monte-Carlo simulation). RESULTS: The incremental cost effectiveness ratio (ICER) of treatment was euro2369 per quality adjusted life year (QALY). In univariate sensitivity analysis the ICER was most sensitive to uncertainty concerning the utility of unilateral visual impairment-for example, if this utility was 0.99, the ICER would be euro9148/QALY. Monte-Carlo simulation yielded a 95% uncertainty interval for the ICER of euro710/QALY to euro38 696/QALY; the probability of an ICER smaller than euro20 000/QALY was 95%. CONCLUSION: Treatment for amblyopia is likely to be very cost effective. Much of the uncertainty in results comes from the uncertainty regarding the effect of amblyopia on quality of life. In order to reduce this uncertainty the impact of amblyopia on utility should be investigated.

Amblyopia↗

A hierarchical mixture of Markov models for finding biologically active metabolic paths using gene expression and protein classes.

With the recent development of experimental high-throughput techniques, the type and volume of accumulating biological data have extremely increased these few years. Mining from different types of data might lead us to find new biological insights. We present a new methodology for systematically combining three different datasets to find biologically active metabolic paths/patterns. This method consists of two steps: First it synthesizes metabolic paths from a given set of chemical reactions, which are already known and whose enzymes are co-expressed, in an efficient manner. It then represents the obtained metabolic paths in a more comprehensible way through estimating parameters of a probabilistic model by using these synthesized paths. This model is built upon an assumption that an entire set of chemical reactions corresponds to a Markov state transition diagram. Furthermore, this model is a hierarchical latent variable model, containing a set of protein classes as a latent variable, for clustering input paths in terms of existing knowledge of protein classes. We tested the performance of our method using a main pathway of glycolysis, and found that our method achieved higher predictive performance for the issue of classifying gene expressions than those obtained by other unsupervised methods. We further analyzed the estimated parameters of our probabilistic models, and found that biologically active paths were clustered into only two or three patterns for each expression experiment type, and each pattern suggested some new long-range relations in the glycolysis pathway.

Computer Simulation↗

A model-based approach to capture genetic variation for future association studies.

Genome-wide association studies are still constrained by the cost of genotyping. For this reason, the selection of a reduced set of markers or tags able to capture a significant proportion of the genetic variation is an important aspect of these studies. Most tagging SNP selection methods have been successful in capturing the genetic variation of the data from which the tags have been chosen. However, when these tags are used in an independent data set, a significant proportion of the remaining SNPs (non-tags) are not captured and, in most cases, there is no information on which SNPs are captured. We propose to use a probabilistic model to predict the non-tags based on a set of tags, as a way to capture genetic variation. An important advantage of this method is that it directly predicts the genotype of the non-tags with which we can test for association with the phenotype and which could help to elucidate the location of genes responsible for increasing disease susceptibility. Additionally, this method provides an estimate of the probabilities with which the predictions are made, which reflects the confidence of the probabilistic model. We also propose new methods to select the tagging SNPs. We empirically show by using HapMap data that our approach is able to capture significantly more genetic variation than methods based solely on a pairwise LD measure.

Algorithms↗

Deficiency in POLE Exonuclease Causes Synthetic Lethality in Highly Aneuploid Cancer Cells.

UNLABELLED: Aneuploidy is a hallmark of cancer and is associated with drug resistance and poor clinical outcomes across diverse cancer types. However, no therapies have been clinically established to target highly aneuploid tumors. By analyzing nearly half a million tumor samples subjected to comprehensive genomic profiling, we identified a striking mutual exclusivity between POLE exonuclease domain mutations and high aneuploidy burden. This observation was independently validated using data from The Cancer Genome Atlas (TCGA) and the Cancer Cell Line Encyclopedia (CCLE). Probabilistic modeling revealed that the elevated quantity and unique spectrum of mutations induced by POLE exonuclease deficiency increase the likelihood of inactivating essential genes on chromosome arms harboring losses, leading to a synthetic lethal phenotype in highly aneuploid cells. Functional experiments demonstrated that POLE exonuclease activity is essential for the viability of highly aneuploid cancer cell lines but dispensable in diploid cells. These findings suggest that selective inhibition of POLE exonuclease activity may represent a promising therapeutic strategy for targeting highly aneuploid tumors. SIGNIFICANCE: An integrated approach using large-scale genomic analyses, probabilistic modeling and functional validation identified POLE exonuclease as a potential synthetic lethal target to overcome cancer aneuploidy.

Humans↗