Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian networks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Estimation of distribution algorithms with Kikuchi approximations.

The question of finding feasible ways for estimating probability distributions is one of the main challenges for Estimation of Distribution Algorithms (EDAs). To estimate the distribution of the selected solutions, EDAs use factorizations constructed according to graphical models. The class of factorizations that can be obtained from these probability models is highly constrained. Expanding the class of factorizations that could be employed for probability approximation is a necessary step for the conception of more robust EDAs. In this paper we introduce a method for learning a more general class of probability factorizations. The method combines a reformulation of a probability approximation procedure known in statistical physics as the Kikuchi approximation of energy, with a novel approach for finding graph decompositions. We present the Markov Network Estimation of Distribution Algorithm (MN-EDA), an EDA that uses Kikuchi approximations to estimate the distribution, and Gibbs Sampling (GS) to generate new points. A systematic empirical evaluation of MN-EDA is done in comparison with different Bayesian network based EDAs. From our experiments we conclude that the algorithm can outperform other EDAs that use traditional methods of probability approximation in the optimization of functions with strong interactions among their variables.

Algorithms↗

Increasing feasibility of optimal gene network estimation.

Disentangling networks of regulation of gene expression is a major challenge in the field of computational biology. Harvesting the information contained in microarray data sets is a promising approach towards this challenge. We propose an algorithm for the optimal estimation of Bayesian networks from microarray data, which reduces the CPU time and memory consumption of previous algorithms. We prove that the space complexity can be reduced from O(n(2) x 2(n)) to O(2(n)), and that the expected calculation time can be reduced from O(n(2) x 2(n)) to O(n x 2(n)), where n is the number of genes. We make intrinsic use of a limitation of the maximal number of regulators of each gene, which has biological as well as statistical justifications. The improvements are significant for some applications in research.

Algorithms↗

Using graphical models and genomic expression data to statistically validate models of genetic regulatory networks.

We propose a model-driven approach for analyzing genomic expression data that permits genetic regulatory networks to be represented in a biologically interpretable computational form. Our models permit latent variables capturing unobserved factors, describe arbitrarily complex (more than pair-wise) relationships at varying levels of refinement, and can be scored rigorously against observational data. The models that we use are based on Bayesian networks and their extensions. As a demonstration of this approach, we utilize 52 genomes worth of Affymetrix GeneChip expression data to correctly differentiate between alternative hypotheses of the galactose regulatory network in S. cerevisiae. When we extend the graph semantics to permit annotated edges, we are able to score models describing relationships at a finer degree of specification.

Bayes Theorem↗

Feature selection in Bayesian classifiers for the prognosis of survival of cirrhotic patients treated with TIPS.

The transjugular intrahepatic portosystemic shunt (TIPS) is a treatment for cirrhotic patients with portal hypertension. A subgroup of patients dies in the first 6 months and another subgroup lives a long period of time. Nowadays, no risk factors have been identified in order to determine how long a patient will survive. An empirical study for predicting the survival rate within the first 6 months after TIPS placement is conducted using a clinical database with 107 cases and 77 variables. Applications of Bayesian classification models, based on Bayesian networks, to medical problems have become popular in the last years. Feature subset selection is useful due to the heterogeneity of the medical databases where not all the variables are required to perform the classification. In this paper, filter and wrapper approaches based on the feature subset selection are adapted to induce Bayesian classifiers (naive Bayes, selective naive Bayes, semi naive Bayes, tree augmented naive Bayes, and k-dependence Bayesian classifier) and are applied to distinguish between the two subgroups of cirrhotic patients. The estimated accuracies obtained tally with the results of previous studies. Moreover, the medical significance of the subset of variables selected by the classifiers along with the comprehensibility of Bayesian models is greatly appreciated by physicians.

Bayes Theorem↗

Optimal dose and exercise modality to improve HbA1c in older adults with type 2 diabetes mellitus: a systematic review with pairwise, network, and dose-response meta-analyses.

We aimed to compare exercise modalities and evaluate dose-response relationships with glycemic control including continuous aerobic exercise (CAE), resistance training (RT), combined exercise (CE), mind-body exercise (MBE), and high-intensity interval training (HIIT) in older adults with type 2 diabetes mellitus (T2DM). Three databases were searched for randomized controlled trials of exercise interventions in older adults with T2DM reporting glycated hemoglobin (HbA1c). Pairwise, Bayesian network, and dose-response meta-analyses were conducted. Compared with control, HIIT demonstrated the largest estimated reduction (MD = -0.95%; 95% CrI -1.45, -0.49), followed by CE (MD = -0.59%; 95% CrI -0.93, -0.25), CAE (MD = -0.46%; 95% CrI -0.69, -0.24), MBE (MD = -0.42%; 95% CrI -0.76, -0.10), and RT (MD = -0.29%; 95% CrI -0.51, -0.08). Dose-response network meta-analyses suggested a non-linear association between overall exercise dose and HbA1c reduction, with maximal estimated benefits at approximately 704 METs-min/week with the 95% CrI excluding zero between 241 and 920 METs-min/week. HIIT demonstrated the steepest estimated dose-response relationship, but with wider credible intervals. Other exercise modalities showed more gradual dose-response patterns across their estimated effective ranges. Our findings suggest that exercise prescription for older adults with T2DM should be individualized according to exercise modality, dose, and health status.

Humans↗

Predicting disease outcome of non-invasive transitional cell carcinoma of the urinary bladder using an artificial neural network model: results of patient follow-up for 15 years or longer.

BACKGROUND: Patients with non-invasive (Ta/T1) transitional cell carcinoma (TCC) of the urinary bladder are often observed without progression in the long-term follow-up period, although many of them experience recurrence of disease. It is difficult to accurately predict the disease outcome of each patient with Ta/T1 TCC using conventional prognostic criteria. In this study, we examined the usefulness of artificial neural networks (ANNs) to predict the long-term disease outcome of patients with TCC of the urinary bladder. METHODS: A retrospective, prognostic study of 90 patients with Ta/T1 TCC of the urinary bladder, diagnosed by transurethral resection of the bladder tumor between April 1981 and March 1985, and then followed up for 15 years or longer, was carried out. Data were analyzed using the Bayesian network tool of SPSS Neural Connection 2.1. The input neural data consisted of tumor stage, grade, tumor number, age, gender, tumor architecture and estimates of mean nuclear volume. The data set was randomly divided into 68 training and 22 testing examples for the prediction of disease progression and tumor recurrence within 15 years. RESULTS: During 15 years follow-up, tumor recurrence was noted in 42/90 (47%) Ta/T1 tumors. The ANN model could not predict tumor recurrence. Conversely, disease progression was noted in 17/90 (19%) Ta/T1 tumors, and, in the test set, 4/22 (18%) Ta/T1 tumors underwent disease progression. The sensitivity of the ANN model to predict progression was 100% (specificity 67%; positive predictive value 40%; negative predictive value 100%). Patients who were judged to have a favorable prognosis using ANN analysis did not progress within the 15-year follow-up period. CONCLUSION: The results of the ANN study indicate that long-term progression-free survival of patients with non-invasive TCC of the urinary bladder can be precisely predicted. A favorable prognosis using ANNs would be one of the exclusion criteria for immediate or future total cystectomy.

Adult↗

Modelling regulatory pathways in E. coli from time series expression profiles.

MOTIVATION: Cells continuously reprogram their gene expression network as they move through the cell cycle or sense changes in their environment. In order to understand the regulation of cells, time series expression profiles provide a more complete picture than single time point expression profiles. Few analysis techniques, however, are well suited to modelling such time series data. RESULTS: We describe an approach that naturally handles time series data with the capabilities of modelling causality, feedback loops, and environmental or hidden variables using a Dynamic Bayesian network. We also present a novel way of combining prior biological knowledge and current observations to improve the quality of analysis and to model interactions between sets of genes rather than individual genes. Our approach is evaluated on time series expression data measured in response to physiological changes that affect tryptophan metabolism in E. coli. Results indicate that this approach is capable of finding correlations between sets of related genes.

Adaptation, Physiological↗

Graphical-Model-based Morphometric Analysis.

We propose a novel method for voxel-based morphometry (VBM), which we call Graphical-Model-based Morphometric Analysis (GAMMA), to identify morphological abnormalities automatically, and to find complex probabilistic associations among voxels in magnetic-resonance images and clinical variables. GAMMA is a fully automatic, nonparametric morphometric-analysis algorithm, with high sensitivity and specificity. It uses a Bayesian network to represent the associations among voxels and the function variable, and uses a contextual-clustering method based on a Markov random field to find clusters in which all voxels have similar associations with the function variable. We use loopy belief propagation to infer the unobserved label field and belief map. As opposed to voxel-based morphometric methods based on general linear models, GAMMA is capable of identifying nonlinear associations among the function variable and voxels. Compared with our previous approach, a Bayesian morphometry algorithm, GAMMA has greater sensitivity, specificity, and computational efficiency.

Algorithms↗

Filtering high-throughput protein-protein interaction data using a combination of genomic features.

BACKGROUND: Protein-protein interaction data used in the creation or prediction of molecular networks is usually obtained from large scale or high-throughput experiments. This experimental data is liable to contain a large number of spurious interactions. Hence, there is a need to validate the interactions and filter out the incorrect data before using them in prediction studies. RESULTS: In this study, we use a combination of 3 genomic features -- structurally known interacting Pfam domains, Gene Ontology annotations and sequence homology -- as a means to assign reliability to the protein-protein interactions in Saccharomyces cerevisiae determined by high-throughput experiments. Using Bayesian network approaches, we show that protein-protein interactions from high-throughput data supported by one or more genomic features have a higher likelihood ratio and hence are more likely to be real interactions. Our method has a high sensitivity (90%) and good specificity (63%). We show that 56% of the interactions from high-throughput experiments in Saccharomyces cerevisiae have high reliability. We use the method to estimate the number of true interactions in the high-throughput protein-protein interaction data sets in Caenorhabditis elegans, Drosophila melanogaster and Homo sapiens to be 27%, 18% and 68% respectively. Our results are available for searching and downloading at http://helix.protein.osaka-u.ac.jp/htp/. CONCLUSION: A combination of genomic features that include sequence, structure and annotation information is a good predictor of true interactions in large and noisy high-throughput data sets. The method has a very high sensitivity and good specificity and can be used to assign a likelihood ratio, corresponding to the reliability, to each interaction.

Animals↗

Use of multiple data streams to conduct Bayesian biologic surveillance.

INTRODUCTION: Emergency department (ED) records and over-the-counter (OTC) sales data are two of the most commonly used sources of data for syndromic surveillance. The majority of detection algorithms monitor these data sources separately and either do not combine them or combine them in an ad hoc fashion. This report outlines a new causal model that combines the two data sources coherently to perform outbreak detection. OBJECTIVES: This report describes the extension of the Population-wide Anomaly Detection and Assessment (PANDA) Bayesian biologic surveillance algorithm to combine information from multiple data streams. It also outlines the assumptions and techniques used to make this approach scalable for real-time surveillance of a large population. METHODS: A causal Bayesian network model used previously was extended to incorporate evidence from daily OTC sales data. At the level of individual persons, the actions that result in the purchase of OTC products and in admission to an ED were modeled. RESULTS: Preliminary results indicate that this model has a tractable running time consisting of 209 seconds for initialization and approximately 4 seconds for every hour's worth of ED data, as measured on a Pentium-4 three-Gigahertz machine with two Gigabytes of RAM. CONCLUSION: Preliminary results for surveillance using a new Bayesian algorithm that models the interaction between ED and OTC data are positive regarding the run time of the algorithm.

Algorithms↗

FINEX: a Probabilistic Expert System for forensic identification.

A series of recent papers have shown how to formulate complex problems of forensic DNA identification inference, such as occur in disputed paternity or criminal identification cases, in terms of Probabilistic Expert Systems (PESs). However, at the present time, general purpose PES software is not particularly well suited to the repetitive tasks of: specifying an appropriate set of marker networks for a specific problem; for editing the many local conditional probability tables; and combining evidence from several genetic markers to evaluate likelihoods. Here, I describe a user-friendly prototype software tool called FINEX developed both to automate such tasks and also to evaluate likelihoods of interest. Ease of use is achieved by a graphical specification language that enables a user to quickly specify a range of forensic DNA problems. I describe the algorithms by which FINEX converts the user input in the graphical specification language and data on observed markers to the Bayesian networks used in PES.

Algorithms↗

Gene networks as a tool to understand transcriptional regulation.

Gene regulatory networks, or simply gene networks (GNs), have shown to be a promising approach that the bioinformatics community has been developing for studying regulatory mechanisms in biological systems. GNs are built from the genome-wide high-throughput gene expression data that are often available from DNA microarray experiments. Conceptually, GNs are (un)directed graphs, where the nodes correspond to the genes and a link between a pair of genes denotes a regulatory interaction that occurs at transcriptional level. In the present study, we had two objectives: 1) to develop a framework for GN reconstruction based on a Bayesian network model that captures direct interactions between genes through nonparametric regression with B-splines, and 2) to demonstrate the potential of GNs in the analysis of expression data of a real biological system, the yeast pheromone response pathway. Our framework also included a number of search schemes to learn the network. We present an intuitive notion of GN theory as well as the detailed mathematical foundations of the model. A comprehensive analysis of the consistency of the model when tested with biological data was done through the analysis of the GNs inferred for the yeast pheromone pathway. Our results agree fairly well with what was expected based on the literature, and we developed some hypotheses about this system. Using this analysis, we intended to provide a guide on how GNs can be effectively used to study transcriptional regulation. We also discussed the limitations of GNs and the future direction of network analysis for genomic data. The software is available upon request.

Bayes Theorem↗

Prediction of protein solvent accessibility using support vector machines.

A Support Vector Machine learning system has been trained to predict protein solvent accessibility from the primary structure. Different kernel functions and sliding window sizes have been explored to find how they affect the prediction performance. Using a cut-off threshold of 15% that splits the dataset evenly (an equal number of exposed and buried residues), this method was able to achieve a prediction accuracy of 70.1% for single sequence input and 73.9% for multiple alignment sequence input, respectively. The prediction of three and more states of solvent accessibility was also studied and compared with other methods. The prediction accuracies are better than, or comparable to, those obtained by other methods such as neural networks, Bayesian classification, multiple linear regression, and information theory. In addition, our results further suggest that this system may be combined with other prediction methods to achieve more reliable results, and that the Support Vector Machine method is a very useful tool for biological sequence analysis.

Bayes Theorem↗

A probabilistic rule-based expert system.

This paper explores a medical expert system combining techniques of Bayesian network modelling with ideas of weighted inference rules. The weights of the individual rules can be estimated objectively from a training set of actual cases; and they can be used in a Monte Carlo stimulation to estimate objectively conditional probabilities of diagnosis given particular combinations of symptoms. The paper describes and evaluates a medical expert system built according to this design. The diagnostic accuracy of the program was found to be similar to that obtained through the usual application of Bayes theorem with the assumption of conditional independence of symptoms given disease, even though the Bayesian classifier has more than 70 times as many numerical parameters. The method may be promising in cases where small training sets do not permit accurate estimation of large numbers of parameters.

Abdominal Pain↗

Understanding tuberculosis epidemiology using structured statistical models.

Molecular epidemiological studies can provide novel insights into the transmission of infectious diseases such as tuberculosis. Typically, risk factors for transmission are identified using traditional hypothesis-driven statistical methods such as logistic regression. However, limitations become apparent in these approaches as the scope of these studies expand to include additional epidemiological and bacterial genomic data. Here we examine the use of Bayesian models to analyze tuberculosis epidemiology. We begin by exploring the use of Bayesian networks (BNs) to identify the distribution of tuberculosis patient attributes (including demographic and clinical attributes). Using existing algorithms for constructing BNs from observational data, we learned a BN from data about tuberculosis patients collected in San Francisco from 1991 to 1999. We verified that the resulting probabilistic models did in fact capture known statistical relationships. Next, we examine the use of newly introduced methods for representing and automatically constructing probabilistic models in structured domains. We use statistical relational models (SRMs) to model distributions over relational domains. SRMs are ideally suited to richly structured epidemiological data. We use a data-driven method to construct a statistical relational model directly from data stored in a relational database. The resulting model reveals the relationships between variables in the data and describes their distribution. We applied this procedure to the data on tuberculosis patients in San Francisco from 1991 to 1999, their Mycobacterium tuberculosis strains, and data on contact investigations. The resulting statistical relational model corroborated previously reported findings and revealed several novel associations. These models illustrate the potential for this approach to reveal relationships within richly structured data that may not be apparent using conventional statistical approaches. We show that Bayesian methods, in particular statistical relational models, are an important tool for understanding infectious disease epidemiology.

Adult↗

Adaptive diagnosis in distributed systems.

Real-time problem diagnosis in large distributed computer systems and networks is a challenging task that requires fast and accurate inferences from potentially huge data volumes. In this paper, we propose a cost-efficient, adaptive diagnostic technique called active probing. Probes are end-to-end test transactions that collect information about the performance of a distributed system. Active probing uses probabilistic reasoning techniques combined with information-theoretic approach, and allows a fast online inference about the current system state via active selection of only a small number of most-informative tests. We demonstrate empirically that the active probing scheme greatly reduces both the number of probes (from 60% to 75% in most of our real-life applications), and the time needed for localizing the problem when compared with nonadaptive (preplanned) probing schemes. We also provide some theoretical results on the complexity of probe selection, and the effect of "noisy" probes on the accuracy of diagnosis. Finally, we discuss how to model the system's dynamics using dynamic Bayesian networks (DBNs), and an efficient approximate approach called sequential multifault; empirical results demonstrate clear advantage of such approaches over "static" techniques that do not handle system's changes.

Algorithms↗

Quality control in nerve conduction studies with coupled knowledge-based system approach.

Contemporary equipment used for nerve conduction studies is usually capable of computerized measurement of latency, amplitude, duration, and area of nerve and muscle action potentials and resulting conduction velocities. Abnormalities can be due to technical error or disease. Identification of technical error is a major element of quality control in electromyography, and artificial intelligence could be useful for this purpose. We have developed a coupled knowledge-based prototype system (QUALICON) to assess the correctness of recording and stimulating characteristics in routine conduction studies. QUALICON extracts numeric features from CMAPs or SNAPs, which are translated into symbolic form to drive a Bayesian network. The network uses high-level knowledge to infer the quality of stimulating and recording electrode placement as well as polarity and stimulus strength making recommendations as to the likely technical error when abnormal potentials are detected. A preliminary assessment shows that QUALICON performs as well as manual assessment performed by professionals.

Action Potentials↗

Discovering structural correlations in alpha-helices.

We have developed a new representation for structural and functional motifs in protein sequences based on correlations between pairs of amino acids and applied it to alpha-helical and beta-sheet sequences. Existing probabilistic methods for representing and analyzing protein sequences have traditionally assumed conditional independence of evidence. In other words, amino acids are assumed to have no effect on each other. However, analyses of protein structures have repeatedly demonstrated the importance of interactions between amino acids in conferring both structure and function. Using Bayesian networks, we are able to model the relationships between amino acids at distinct positions in a protein sequence in addition to the amino acid distributions at each position. We have also developed an automated program for discovering sequence correlations using standard statistical tests and validation techniques. In this paper, we test this program on sequences from secondary structure motifs, namely alpha-helices and beta-sheets. In each case, the correlations our program discovers correspond well with known physical and chemical interactions between amino acids in structures. Furthermore, we show that, using different chemical alphabets for the amino acids, we discover structural relationships based on the same chemical principle used in constructing the alphabet. This new representation of 3-dimensional features in protein motifs, such as those arising from structural or functional constraints on the sequence, can be used to improve sequence analysis tools including pattern analysis and database search.

Amino Acid Sequence↗