Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “probabilistic modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Protein classification using probabilistic chain graphs and the Gene Ontology structure.

MOTIVATION: Probabilistic graphical models have been developed in the past for the task of protein classification. In many cases, classifications obtained from the Gene Ontology have been used to validate these models. In this work we directly incorporate the structure of the Gene Ontology into the graphical representation for protein classification. We present a method in which each protein is represented by a replicate of the Gene Ontology structure, effectively modeling each protein in its own 'annotation space'. Proteins are also connected to one another according to different measures of functional similarity, after which belief propagation is run to make predictions at all ontology terms. RESULTS: The proposed method was evaluated on a set of 4879 proteins from the Saccharomyces Genome Database whose interactions were also recorded in the GRID project. Results indicate that direct utilization of the Gene Ontology improves predictive ability, outperforming traditional models that do not take advantage of dependencies among functional terms. Average increase in accuracy (precision) of positive and negative term predictions of 27.8% (2.0%) over three different similarity measures and three subontologies was observed. AVAILABILITY: C/C++/Perl implementation is available from authors upon request.

Algorithms↗

Indelign: a probabilistic framework for annotation of insertions and deletions in a multiple alignment.

MOTIVATION: A quantitative study of molecular evolutionary events such as substitutions, insertions and deletions from closely related genomes requires (1) an accurate multiple sequence alignment program and (2) a method to annotate the insertions and deletions that explain the 'gaps' in the alignment. Although the former requirement has been extensively addressed, the latter problem has received little attention, especially in a comprehensive probabilistic framework. RESULTS: Here, we present Indelign, a program that uses a probabilistic evolutionary model to compute the most likely scenario of insertions and deletions consistent with an input multiple alignment. It is also capable of modifying the given alignment so as to obtain a better agreement with the evolutionary model. We find close to optimal performance and substantial improvement over alternative methods, in tests of Indelign on synthetic data. We use Indelign to analyze regulatory sequences in Drosophila, and find an excess of insertions over deletions, which is different from what has been reported for neutral sequences. AVAILABILITY: The Indelign program may be downloaded from the website http://veda.cs.uiuc.edu/indelign/ SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.

Algorithms↗

A divide and conquer strategy for recapitulating whole genome 3D structure using Hi-C data.

The three dimensional (3D) spatial organization of the genome is closely linked to biological functions and can be captured by Hi-C assays through interrogating genome-wide chromatin interactions. Methodologies for inferring 3D structures from Hi-C data summarized as a two-dimensional (2D) contact matrix can be broadly placed within the paradigms of optimization-based and sampling-based. Many optimization-based methods are capable of constructing whole genome 3D structures but do not account for spatial dependency in the 2D data matrix nor cell heterogeneity in bulk Hi-C data, which provide an average over millions of cells. Sampling-based methods, on the other hand, are probabilistic model-based and can account for not only dependency, heterogeneity, but also other features inherent in Hi-C data, such as over-dispersion and sparsity. However, whole-genome 3D structure recapitulation is too computationally expensive for sampling-based methods, while chromosome-by-chromosome strategies for sampling-based methods ignore important information on inter-chromosomal contacts. To address these issues, we propose the truncated Random effect EXpression-cut and paste (tREX-cap) method, which applies the tREX model within a divide and conquer strategy. The resulting method inherits the good data-feature-cognizant properties of tREX and, in the meantime, can efficiently infer the whole genome 3D structure. We demonstrate the performance of tREX-cap through an extensive simulation study and analyses of a Hi-C lymphoblastoid dataset and a Hi-C IMR90 dataset.

Humans↗

INCLUSive: A web portal and service registry for microarray and regulatory sequence analysis.

INCLUSive is a suite of algorithms and tools for the analysis of gene expression data and the discovery of cis-regulatory sequence elements. The tools allow normalization, filtering and clustering of microarray data, functional scoring of gene clusters, sequence retrieval, and detection of known and unknown regulatory elements using probabilistic sequence models and Gibbs sampling. All tools are available via different web pages and as web services. The web pages are connected and integrated to reflect a methodology and facilitate complex analysis using different tools. The web services can be invoked using standard SOAP messaging. Example clients are available for download to invoke the services from a remote computer or to be integrated with other applications. All services are catalogued and described in a web service registry. The INCLUSive web portal is available for academic purposes at http://www.esat.kuleuven.ac.be/inclusive.

Algorithms↗

NestedMICA: sensitive inference of over-represented motifs in nucleic acid sequence.

NestedMICA is a new, scalable, pattern-discovery system for finding transcription factor binding sites and similar motifs in biological sequences. Like several previous methods, NestedMICA tackles this problem by optimizing a probabilistic mixture model to fit a set of sequences. However, the use of a newly developed inference strategy called Nested Sampling means NestedMICA is able to find optimal solutions without the need for a problematic initialization or seeding step. We investigate the performance of NestedMICA in a range scenario, on synthetic data and a well-characterized set of muscle regulatory regions, and compare it with the popular MEME program. We show that the new method is significantly more sensitive than MEME: in one case, it successfully extracted a target motif from background sequence four times longer than could be handled by the existing program. It also performs robustly on synthetic sequences containing multiple significant motifs. When tested on a real set of regulatory sequences, NestedMICA produced motifs which were good predictors for all five abundant classes of annotated binding sites.

Base Sequence↗

ProMiR II: a web server for the probabilistic prediction of clustered, nonclustered, conserved and nonconserved microRNAs.

ProMiR is a web-based service for the prediction of potential microRNAs (miRNAs) in a query sequence of 60-150 nt, using a probabilistic colearning model. Identification of miRNAs requires a computational method to predict clustered and nonclustered, conserved and nonconserved miRNAs in various species. Here we present an improved version of ProMiR for identifying new clusters near known or unknown miRNAs. This new version, ProMiR II, integrates additional evidence, such as free energy data, G/C ratio, conservation score and entropy of candidate sequences, for more controllable prediction of miRNAs in mouse and human genomes. It also provides a wider range of services, e.g. the prediction of miRNA genes in long nonrelated sequences such as viral genomes. Importantly, we have validated this method using several case studies. All data used in ProMiR II are structured in the MySQL database for efficient analysis. The ProMiR II web server is available at http://cbit.snu.ac.kr/~ProMiR2/.

Animals↗

Assessment of negative and positive symptoms in schizophrenia.

Reliable, convenient rating scales to assess negative and positive symptoms in schizophrenia are necessary to evaluate further the theoretical and clinical importance of this division of symptoms and signs. The authors describe the application of the Rasch model, a probabilistic, item-independent, and sample-independent test construction procedure to the development of scales for both types of symptoms. The scales for negative and positive symptoms, which are based separately on the Schedule for Affective Disorders and Schizophrenia-Current (SADS-C) and the Nurses' Observation Scale for Inpatient Evaluation (NOSIE), demonstrated excellent reliability and temporal stability (i.e., yielded a rank order of patients that remained stable over time). The pattern of interscale correlations supports the view that positive symptoms, cognitive-affective negative symptoms, and social withdrawal are independent of one another.

Adult↗

The cost-effectiveness of elective Cesarean delivery for HIV-infected women with detectable HIV RNA during pregnancy.

OBJECTIVES: To determine the net health consequences, costs, and cost-effectiveness of alternative delivery strategies for HIV-infected pregnant women with detectable HIV RNA in the USA. DESIGN: Cost-effectiveness analysis using a probabilistic decision model. METHODS: The model compared two strategies: elective Cesarean section and vaginal delivery. Data for HIV transmission rate, maternal death rate, health-related quality of life and costs were obtained from the literature, national databases, and a tertiary hospital's cost accounting system. Model outcomes included total lifetime costs, quality-adjusted life expectancy, maternal death rate, HIV transmission rate, and incremental cost-effectiveness ratios. RESULTS: Elective Cesarean section resulted in a vertical HIV transmission rate of 34.9 per 1000 births compared with 62.3 per 1000 births for vaginal delivery. Elective Cesarean section was more effective (38.7 quality adjusted life years per mother and child pair) and less costly ($10600 per delivery) than trial of labor (38.2 combined quality adjusted life years at a cost of $14500 per delivery). However, elective Cesarean section increased maternal mortality by 2.4 deaths per 100000 deliveries. The results were consistent over a wide range of the variables, but were sensitive to the risk of HIV transmission with vaginal delivery and the relative risk of HIV transmission with elective Cesarean section. CONCLUSIONS: In pregnant HIV-infected women with detectable HIV RNA, elective Cesarean section would reduce total costs and increase overall quality-adjusted life expectancy for the mother-child pair, albeit at a slight loss of quality adjusted life expectancy to the mother.

Adult↗

Decline in HIV infectivity following the introduction of highly active antiretroviral therapy.

OBJECTIVE: Little is known about the degree to which widespread use of antiretroviral therapy in a community reduces uninfected individuals' risk of acquiring HIV. We estimated the degree to which the probability of HIV infection from an infected partner (the infectivity) declined following the introduction of highly active antiretroviral therapy (HAART) in San Francisco. DESIGN: Homosexual men from the San Francisco Young Men's Health Study, who were initially uninfected with HIV, were asked about sexual practices, and tested for HIV antibodies at each of four follow-up visits during a 6-year period spanning the advent of widespread use of HAART (1994-1999). METHODS: We estimated the infectivity of HIV (per-partnership probability of transmission from an infected partner) using a probabilistic risk model based on observed incident infections and self-reported sexual risk behavior, and tested the hypothesis that infectivity was the same before and after HAART was introduced. RESULTS: A total of 534 homosexual men were evaluated. Decreasing trends in HIV seroincidence were observed despite increases in reported number of unprotected receptive anal intercourse partners. Conservatively assuming a constant prevalence of HIV infection between 1994 and 1999, HIV infectivity decreased from 0.120 prior to widespread use of HAART, to 0.048 after the widespread use of HAART- a decline of 60% (P=0.028). CONCLUSIONS: Use of HAART by infected persons in a community appears to reduce their infectiousness and therefore may provide an important HIV prevention tool.

Adolescent↗

The cost-effectiveness of elective Cesarean delivery to prevent hepatitis C transmission in HIV-coinfected women.

OBJECTIVES: To determine the net health consequences, costs, and cost-effectiveness of elective Cesarean delivery (C-section) to prevent perinatal transmission of hepatitis C virus (HCV) in HIV/HCV-coinfected women with suppressed HIV RNA but detectable HCV RNA. DESIGN: Cost-effectiveness analysis using a probabilistic decision model. METHODS: The model compared two strategies: (i) C-section for all coinfected women with suppressed HIV RNA but detectable HCV RNA; (ii) C-section only when indicated based on fetal status. Outcomes included vertical transmission of HCV, maternal mortality, quality-adjusted life expectancy, delivery and HCV treatment costs, and incremental cost-effectiveness ratios. Data were obtained from the literature and national databases. Delivery cost data were from a hospital consortium database. Probability distributions were derived from published confidence intervals or estimated ranges, or calculated using reported sample sizes. RESULTS: Elective C-section in coinfected women with suppressed HIV RNA but detectable HCV RNA would avoid 45 vertical HCV transmissions per 1000 deliveries and increase maternal mortality by one death per 100 000 deliveries. The incremental cost-effectiveness ratio of a recommendation for C-section versus current practice was 3900-6100 dollars per quality-adjusted life year for the mother-child pair. Results are sensitive to the efficacy of C-section in preventing transmission, the probability of vaginal delivery without a recommendation, and rates of maternal acceptance of the recommendation. CONCLUSIONS: Assuming 2000 births/year among HIV/HCV-coinfected women in the United States, a recommendation for elective C-section in these women could avoid an additional 90 perinatal HCV transmissions per year with a risk of one maternal death in 50 years.

Cesarean Section↗

Sheltering--a protective measure following an accidental atmospheric release from a nuclear power plant.

The effectiveness of sheltering the population for reducing radiological effects following an accidental release of radioactivity at a nuclear power plant was investigated. Different levels of respiratory protection and the administration of a thyroid blocking agent were also studied as possible complements to sheltering. Specific conditions were assumed, concerning the high protection factors of regular buildings and the high availability of civil defense shelters. Computations were performed by means of a probabilistic consequence model, which allows a comprehensive description of exposure modes and processes dealing with the implementation of sheltering and which takes into account a broad range of radiological effects. Sheltering, even in regular buildings, was found to be efficient in reducing early fatalities and other non-stochastic effects. However, it was shown that respiratory protection is also needed in order to alleviate stochastic effects and that, for this purpose, expedient individual filtration methods may be satisfactory. Under the conditions studied, sheltering was found to be preferable in most cases over evacuation, as the main immediate protective measure, unless evacuation can be carried out before the radioactive cloud reaches the populated area.

Accidents↗

Natural disasters and the challenge of extreme events: risk management from an insurance perspective.

Loss statistics for natural disasters demonstrate, also after correction for inflation, a dramatic increase of the loss burden since 1950. This increase is driven by a concentration of population and values in urban areas, the development of highly exposed coastal and valley regions, the complexity of modern societies and technologies and probably, also by the beginning consequences of global warming. This process will continue unless remedial action will be taken. Managing the risk from natural disasters starts with identification of the hazards. The next step is the evaluation of the risk, where risk is a function of hazard, exposed values or human lives and the vulnerability of the exposed objects. Probabilistic computer models have been developed for the proper assessment of risks since the late 1980s. The final steps are controlling and financing future losses. Natural disaster insurance plays a key role in this context, but also private parties and governments have to share a part of the risk. A main responsibility of governments is to formulate regulations for building construction and land use. The insurance sector and the state have to act together in order to create incentives for building and business owners to take loss prevention measures. A further challenge for the insurance sector is to transfer a portion of the risk to the capital markets, and to serve better the needs of the poor. Catastrophe bonds and microinsurance are the answer to such challenges. The mechanisms described above have been developed to cope with well-known disasters like earthquakes, windstorms and floods. They can be applied, in principle, also to less well investigated and less frequent extreme disasters: submarine slides, great volcanic eruptions, meteorite impacts and tsunamis which may arise from all these hazards. But there is an urgent need to improve the state of knowledge on these more exotic hazards in order to reduce the high uncertainty in actual risk evaluation to an acceptable level. Due to the rarity of such extreme events, specific risk prevention measures are hardly justified with exception of attempts to divert earth-orbit crossing meteorites from their dangerous path. For the industry it is particularly important to achieve full transparency as regards covered and non-covered risks and to define in a systematic manner the limits of insurability for super-disasters.

Disaster Planning↗

Evidence for recombination in Crimean-Congo hemorrhagic fever virus.

Crimean-Congo hemorrhagic fever (CCHF) virus has attracted considerable attention recently and a number of phylogenetic studies have been published, based mostly on partial sequences of S and M RNA segments. In this study, available full-length S, M and L segment sequences of CCHF virus were checked for recombination. Similarity plots and bootscan analysis of the S segment suggested multiple recombination events between southern European, Asian and African CCHF virus strains, with additional evidence provided by phylogenetic trees, the hidden Markov model and probabilistic divergence measures methods. No unambiguous signs of recombination were observed for M and L segments; however, the results did not exclude the possibility of this. These findings, coupled with a recent report on reassortment in CCHF virus, suggest caution when assessing CCHF virus phylogeny based on short sequence fragments.

Animals↗

Predicting protein complex membership using probabilistic network reliability.

Evidence for specific protein-protein interactions is increasingly available from both small- and large-scale studies, and can be viewed as a network. It has previously been noted that errors are frequent among large-scale studies, and that error frequency depends on the large-scale method used. Despite knowledge of the error-prone nature of interaction evidence, edges (connections) in this network are typically viewed as either present or absent. However, use of a probabilistic network that considers quantity and quality of supporting evidence should improve inference derived from protein networks. Here we demonstrate inference of membership in a partially known protein complex by using a probabilistic network model and an algorithm previously used to evaluate reliability in communication networks.

Fungal Proteins↗

Automatic classification of sub-microlitre protein-crystallization trials in 1536-well plates.

A technique for automatically evaluating microbatch (400 nl) protein-crystallization trials is described. This method addresses analysis problems introduced at the sub-microlitre scale, including non-uniform lighting and irregular droplet boundaries. The droplet is segmented from the well using a loopy probabilistic graphical model with a two-layered grid topology. A vector of 23 features is extracted from the droplet image using the Radon transform for straight-edge features and a bank of correlation filters for microcrystalline features. Image classification is achieved by linear discriminant analysis of its feature vector. The results of the automatic method are compared with those of a human expert on 32 1536-well plates. Using the human-labeled images as ground truth, this method classifies images with 85% accuracy and a ROC score of 0.84. This result compares well with the experimental repeatability rate, assessed at 87%. Images falsely classified as crystal-positive variously contain speckled precipitate resembling microcrystals, skin effects or genuine crystals falsely labeled by the human expert. Many images falsely classified as crystal-negative variously contain very fine crystal features or dendrites lacking straight edges. Characterization of these misclassifications suggests directions for improving the method.

Aldose-Ketose Isomerases↗

A functional-dependencies-based Bayesian networks learning method and its application in a mobile commerce system.

This paper presents a new method for learning Bayesian networks from functional dependencies (FD) and third normal form (3NF) tables in relational databases. The method sets up a linkage between the theory of relational databases and probabilistic reasoning models, which is interesting and useful especially when data are incomplete and inaccurate. The effectiveness and practicability of the proposed method is demonstrated by its implementation in a mobile commerce system.

Algorithms↗

Fusion of intelligence information: a Bayesian approach.

The attack that occurred on September 11, 2001 was, in the end, the result of a failure to detect and prevent the terrorist operations that hit the United States. The U.S. government thus faces at this time the daunting tasks of first, drastically increasing its ability to obtain and interpret different types of signals of impending terrorist attacks with sufficient lead time and accuracy, and second, improving its ability to react effectively. One of the main challenges is the fusion of information, from different sources (U.S. or foreign), and of different types (electronic signals, human intelligence. etc.). Fusion thus involves two very distinct and separate issues: communications, i.e., ensuring that the different U.S. and foreign intelligence agencies communicate all relevant and accurate information in a timely fashion and, perhaps more difficult, merging the content of signals, some "sharp" and some "fuzzy," some dependent and some independent into useful information. The focus of this article is on the latter issue, and on the use of the results. In this article, I present a classic probabilistic Bayesian model sometimes used in engineering risk analysis, which can be helpful in the fusion of information because it allows computation of the posterior probability of an event given its prior probability (before the signal is observed) and the quality of the signal characterized by the probabilities of false positive and false negative. Experience suggests that the nature of these errors has been sometimes misunderstood; therefore, I discuss the validity of several possible definitions.

Journal Article↗

Establishing the cost-effectiveness of new pharmaceuticals under conditions of uncertainty--when is there sufficient evidence?

Decisions about which health-care interventions represent adequate value to collectively funded health-care systems are as widespread as they are unavoidable. In the case of new pharmaceuticals, many countries now require formal cost-effectiveness analysis to inform this decision-making process. This requires evidence on parameters associated with health-related utilities, treatment effects, resource use, and costs, for which data from available regulatory trials are invariably absent or highly uncertain. This uncertainty results from a number of factors including the predominance of intermediate end points in the clinical evidence-base and the limited period of follow-up of patients in clinical studies. Despite these imperfections in the evidence base, decisions about whether new pharmaceuticals are sufficiently cost-effective for reimbursement cannot be side-stepped. Data limitations do, however, require the use of rigorous analytical methods to support decision making. Probabilistic decision models and value of information analysis offer a means of structuring decision problems, synthesizing all available data, characterizing the uncertainty in the decision, quantifying the cost of uncertainty, and establishing the expected value of perfect information. This analytical framework is important because it addresses two fundamental questions about new pharmaceuticals. First, is the product expected to be cost-effective on the basis of existing evidence? Second, is additional research concerning the product itself cost-effective? In addressing these questions, the analytical framework can establish when sufficient evidence exists to sustain a claim for a new pharmaceutical to be cost-effective.

Cost-Benefit Analysis↗