Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian networks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Derivation and validation of a Bayesian network to predict pretest probability of venous thromboembolism.

STUDY OBJECTIVE: A Bayesian network can estimate a numeric pretest probability of venous thromboembolism on the basis of values of clinical variables. We determine the accuracy with which a Bayesian network can identify patients with a low pretest probability of venous thromboembolism, defined as less than or equal to 2%. METHODS: Using commercial software, we derived a population of Bayesian networks from 25 input variables collected on 3,145 emergency department (ED) patients with suspected venous thromboembolism who underwent standardized testing, including pulmonary vascular imaging, and 90-day follow-up (11.0% of patients were venous thromboembolism positive). The best-fit Bayesian network was selected using a genetic algorithm. The selected Bayesian network was tested in a validation population of 1,423 ED patients prospectively evaluated for venous thromboembolism, including 90-day follow-up (8.0% were venous thromboembolism positive). The Bayesian network probability estimate was normalized to a score of 0% to 100%. RESULTS: Of 1,423 patients in the validation cohort, 711 (50%; 95% confidence interval [CI] 47% to 52%) had a score less than or equal to 2% that predicted a low pretest probability. Of these 711 patients, 700 (98.5%; 95% CI 97.2% to 99.2%) had no venous thromboembolism at follow-up. CONCLUSION: A Bayesian network, derived and independently validated in ED populations, identified half of the validation cohort as having a low pretest probability (< or =2%); 98.5% of these patients were correctly classified by the network.

Adult↗

Construction of a Bayesian network for mammographic diagnosis of breast cancer.

Bayesian networks use the techniques of probability theory to reason under uncertainty, and have become an important formalism for medical decision support systems. We describe the development and validation of a Bayesian network (MammoNet) to assist in mammographic diagnosis of breast cancer. MammoNet integrates five patient-history features, two physical findings, and 15 mammographic features extracted by experienced radiologists to determine the probability of malignancy. We outline the methods and issues in the system's design, implementation, and evaluation. Bayesian networks provide a potentially useful tool for mammographic decision support.

Bayes Theorem↗

Bayesian model averaging of Bayesian network classifiers over multiple node-orders: application to sparse datasets.

Bayesian model averaging (BMA) can resolve the overfitting problem by explicitly incorporating the model uncertainty into the analysis procedure. Hence, it can be used to improve the generalization performance of Bayesian network classifiers. Until now, BMA of Bayesian network classifiers has only been performed in some restricted forms, e.g., the model is averaged given a single node-order, because of its heavy computational burden. However, it can be hard to obtain a good node-order when the available training dataset is sparse. To alleviate this problem, we propose BMA of Bayesian network classifiers over several distinct node-orders obtained using the Markov chain Monte Carlo sampling technique. The proposed method was examined using two synthetic problems and four real-life datasets. First, we show that the proposed method is especially effective when the given dataset is very sparse. The classification accuracy of averaging over multiple node-orders was higher in most cases than that achieved using a single node-order in our experiments. We also present experimental results for test datasets with unobserved variables, where the quality of the averaged node-order is more important. Through these experiments, we show that the difference in classification performance between the cases of multiple node-orders and single node-order is related to the level of noise, confirming the relative benefit of averaging over multiple node-orders for incomplete data. We conclude that BMA of Bayesian network classifiers over multiple node-orders has an apparent advantage when the given dataset is sparse and noisy, despite the method's heavy computational cost.

Algorithms↗

Estimating gene networks from gene expression data by combining Bayesian network model with promoter element detection.

We present a statistical method for estimating gene networks and detecting promoter elements simultaneously. When estimating a network from gene expression data alone, a common problem is that the number of microarrays is limited compared to the number of variables in the network model, making accurate estimation a difficult task. Our method overcomes this problem by integrating the microarray gene expression data and the DNA sequence information into a Bayesian network model. The basic idea of our method is that, if a parent gene is a transcription factor, its children may share a consensus motif in their promoter regions of the DNA sequences. Our method detects consensus motifs based on the structure of the estimated network, then re-estimates the network using the result of the motif detection. We continue this iteration until the network becomes stable. To show the effectiveness of our method, we conducted Monte Carlo simulations and applied our method to Saccharomyces cerevisiae data as a real application.

Bayes Theorem↗

A comparison of Bayesian network learning algorithms from continuous data.

Learning a Bayesian network from data is an important problem in biomedicine for the automatic construction of decision support systems and inference of plausible causal relations. Most Bayesian network learning algorithms require discrete data; however discretization may impact the quality of the learned structure. In this project, we present a comparison of different approaches for learning from continuous data to identify the most promising one and to quantify the impact of discretization in Bayesian network learning.

Algorithms↗

Estimation of genetic networks and functional structures between genes by using Bayesian networks and nonparametric regression.

We propose a new method for constructing genetic network from gene expression data by using Bayesian networks. We use nonparametric regression for capturing nonlinear relationships between genes and derive a new criterion for choosing the network in general situations. In a theoretical sense, our proposed theory and methodology include previous methods based on Bayes approach. We applied the proposed method to the S. cerevisiae cell cycle data and showed the effectiveness of our method by comparing with previous methods.

Bayes Theorem↗

Bayesian network and nonparametric heteroscedastic regression for nonlinear modeling of genetic network.

We propose a new statistical method for constructing a genetic network from microarray gene expression data by using a Bayesian network. An essential point of Bayesian network construction is the estimation of the conditional distribution of each random variable. We consider fitting nonparametric regression models with heterogeneous error variances to the microarray gene expression data to capture the nonlinear structures between genes. Selecting the optimal graph, which gives the best representation of the system among genes, is still a problem to be solved. We theoretically derive a new graph selection criterion from Bayes approach in general situations. The proposed method includes previous methods based on Bayesian networks. We demonstrate the effectiveness of the proposed method through the analysis of Saccharomyces cerevisiae gene expression data newly obtained by disrupting 100 genes.

Bayes Theorem↗

Bayesian network and nonparametric heteroscedastic regression for nonlinear modeling of genetic network.

We propose a new statistical method for constructing genetic network from microarray gene expression data by using a Bayesian network. An essential point of Bayesian network construction is in the estimation of the conditional distribution of each random variable. We consider fitting nonparametric regression models with heterogeneous error variances to the microarray gene expression data to capture the nonlinear structures between genes. A problem still remains to be solved in selecting an optimal graph, which gives the best representation of the system among genes. We theoretically derive a new graph selection criterion from Bayes approach in general situations. The proposed method includes previous methods based on Bayesian networks. We demonstrate the effectiveness of the proposed method through the analysis of Saccharomyces cerevisiae gene expression data newly obtained by disrupting 100 genes.

Artificial Intelligence↗

Comparison of neural network, Bayesian, and multiple stepwise regression-based limited sampling models to estimate area under the curve.

This study compared limited sampling methods (LSM) of estimating area under the plasma concentration versus time curve (AUC) based on a Bayesian regularized neural network, the Bayesian approach, and multiple forward stepwise regression models from selected concentration-time points. Plasma concentration versus time data sets with a linear two-compartmental pharmacokinetic model were simulated. A limited sampling method based on the forward stepwise regression model was developed and validated. Plasma concentration-time points selected by the stepwise regression model were used for neural network and Bayesian evaluation. In addition, 55 plasma concentration-time profiles from two clinical studies were used to develop and compare the predicted AUC(last) for the three approaches. From simulated data sets, mean prediction errors for AUC(last) estimation were 0.00, -5.32, and -6.06 for the neural network, Bayesian approach, and forward stepwise regression LSM, respectively. Mean square errors were 581, 588, and 618, respectively. For clinical data set, model mean prediction errors were 0.00, 3.51, and 3.87, respectively. Model mean square errors were 30.6, 109, and 76, respectively. For both simulated and clinical data sets, the neural network approach to estimate AUC(last) from selected time points was numerically more precise and significantly less biased than the other two methods.

Antiviral Agents↗

A Bayesian network classification methodology for gene expression data.

We present new techniques for the application of a Bayesian network learning framework to the problem of classifying gene expression data. The focus on classification permits us to develop techniques that address in several ways the complexities of learning Bayesian nets. Our classification model reduces the Bayesian network learning problem to the problem of learning multiple subnetworks, each consisting of a class label node and its set of parent genes. We argue that this classification model is more appropriate for the gene expression domain than are other structurally similar Bayesian network classification models, such as Naive Bayes and Tree Augmented Naive Bayes (TAN), because our model is consistent with prior domain experience suggesting that a relatively small number of genes, taken in different combinations, is required to predict most clinical classes of interest. Within this framework, we consider two different approaches to identifying parent sets which are supported by the gene expression observations and any other currently available evidence. One approach employs a simple greedy algorithm to search the universe of all genes; the second approach develops and applies a gene selection algorithm whose results are incorporated as a prior to enable an exhaustive search for parent sets over a restricted universe of genes. Two other significant contributions are the construction of classifiers from multiple, competing Bayesian network hypotheses and algorithmic methods for normalizing and binning gene expression data in the absence of prior expert knowledge. Our classifiers are developed under a cross validation regimen and then validated on corresponding out-of-sample test sets. The classifiers attain a classification rate in excess of 90% on out-of-sample test sets for two publicly available datasets. We present an extensive compilation of results reported in the literature for other classification methods run against these same two datasets. Our results are comparable to, or better than, any we have found reported for these two sets, when a train-test protocol as stringent as ours is followed.

Bayes Theorem↗

A Bayesian network model for protein fold and remote homologue recognition.

MOTIVATION: The Bayesian network approach is a framework which combines graphical representation and probability theory, which includes, as a special case, hidden Markov models. Hidden Markov models trained on amino acid sequence or secondary structure data alone have been shown to have potential for addressing the problem of protein fold and superfamily classification. RESULTS: This paper describes a novel implementation of a Bayesian network which simultaneously learns amino acid sequence, secondary structure and residue accessibility for proteins of known three-dimensional structure. An awareness of the errors inherent in predicted secondary structure may be incorporated into the model by means of a confusion matrix. Training and validation data have been derived for a number of protein superfamilies from the Structural Classification of Proteins (SCOP) database. Cross validation results using posterior probability classification demonstrate that the Bayesian network performs better in classifying proteins of known structural superfamily than a hidden Markov model trained on amino acid sequences alone.

Amino Acid Sequence↗

Bayesian network analysis of signaling networks: a primer.

High-throughput proteomic data can be used to reveal the connectivity of signaling networks and the influences between signaling molecules. We present a primer on the use of Bayesian networks for this task. Bayesian networks have been successfully used to derive causal influences among biological signaling molecules (for example, in the analysis of intracellular multicolor flow cytometry). We discuss ways to automatically derive a Bayesian network model from proteomic data and to interpret the resulting model.

Algorithms↗

Using protein-protein interactions for refining gene networks estimated from microarray data by Bayesian networks.

We propose a statistical method to estimate gene networks from DNA microarray data and protein-protein interactions. Because physical interactions between proteins or multiprotein complexes are likely to regulate biological processes, using only mRNA expression data is not sufficient for estimating a gene network accurately. Our method adds knowledge about protein-protein interactions to the estimation method of gene networks under a Bayesian statistical framework. In the estimated gene network, a protein complex is modeled as a virtual node based on principal component analysis. We show the effectiveness of the proposed method through the analysis of Saccharomyces cerevisiae cell cycle data. The proposed method improves the accuracy of the estimated gene networks, and successfully identifies some biological facts.

Algorithms↗

Drug delivery optimization through Bayesian networks.

This paper describes how Bayesian Networks can be used in combination with compartmental models to plan Recombinant Human Erythropoietin (r-HuEPO) delivery in the treatment of anemia of chronic uremic patients. Past measurements of hematocrit or hemoglobin concentration in a patient during the therapy can be exploited to adjust the parameters of a compartmental model of the erythropoiesis. This adaptive process allows more accurate patient-specific predictions, and hence a more rational dosage planning. We describe a drug delivery optimization protocol, based on our approach. Some results obtained on real data are presented.

Anemia↗

Predicting the effect of missense mutations on protein function: analysis with Bayesian networks.

BACKGROUND: A number of methods that use both protein structural and evolutionary information are available to predict the functional consequences of missense mutations. However, many of these methods break down if either one of the two types of data are missing. Furthermore, there is a lack of rigorous assessment of how important the different factors are to prediction. RESULTS: Here we use Bayesian networks to predict whether or not a missense mutation will affect the function of the protein. Bayesian networks provide a concise representation for inferring models from data, and are known to generalise well to new data. More importantly, they can handle the noisy, incomplete and uncertain nature of biological data. Our Bayesian network achieved comparable performance with previous machine learning methods. The predictive performance of learned model structures was no better than a naïve Bayes classifier. However, analysis of the posterior distribution of model structures allows biologically meaningful interpretation of relationships between the input variables. CONCLUSION: The ability of the Bayesian network to make predictions when only structural or evolutionary data was observed allowed us to conclude that structural information is a significantly better predictor of the functional consequences of a missense mutation than evolutionary information, for the dataset used. Analysis of the posterior distribution of model structures revealed that the top three strongest connections with the class node all involved structural nodes. With this in mind, we derived a simplified Bayesian network that used just these three structural descriptors, with comparable performance to that of an all node network.

Algorithms↗

Using literature and data to learn Bayesian networks as clinical models of ovarian tumors.

Thanks to its increasing availability, electronic literature has become a potential source of information for the development of complex Bayesian networks (BN), when human expertise is missing or data is scarce or contains much noise. This opportunity raises the question of how to integrate information from free-text resources with statistical data in learning Bayesian networks. Firstly, we report on the collection of prior information resources in the ovarian cancer domain, which includes "kernel" annotations of the domain variables. We introduce methods based on the annotations and literature to derive informative pairwise dependency measures, which are derived from the statistical cooccurrence of the names of the variables, from the similarity of the "kernel" descriptions of the variables and from a combined method. We perform wide-scale evaluation of these text-based dependency scores against an expert reference and against data scores (the mutual information (MI) and a Bayesian score). Next, we transform the text-based dependency measures into informative text-based priors for Bayesian network structures. Finally, we report the benefit of such informative text-based priors on the performance of a Bayesian network for the classification of ovarian tumors from clinical data.

Artificial Intelligence↗

Using Bayesian networks to analyze expression data.

DNA hybridization arrays simultaneously measure the expression level for thousands of genes. These measurements provide a "snapshot" of transcription levels within the cell. A major challenge in computational biology is to uncover, from such measurements, gene/protein interactions and key biological features of cellular systems. In this paper, we propose a new framework for discovering interactions between genes based on multiple expression measurements. This framework builds on the use of Bayesian networks for representing statistical dependencies. A Bayesian network is a graph-based model of joint multivariate probability distributions that captures properties of conditional independence between variables. Such models are attractive for their ability to describe complex stochastic processes and because they provide a clear methodology for learning from (noisy) observations. We start by showing how Bayesian networks can describe interactions between genes. We then describe a method for recovering gene interactions from microarray data using tools for learning Bayesian networks. Finally, we demonstrate this method on the S. cerevisiae cell-cycle measurements of Spellman et al. (1998).

Algorithms↗

Predicting the prognosis of breast cancer by integrating clinical and microarray data with Bayesian networks.

MOTIVATION: Clinical data, such as patient history, laboratory analysis, ultrasound parameters--which are the basis of day-to-day clinical decision support--are often underused to guide the clinical management of cancer in the presence of microarray data. We propose a strategy based on Bayesian networks to treat clinical and microarray data on an equal footing. The main advantage of this probabilistic model is that it allows to integrate these data sources in several ways and that it allows to investigate and understand the model structure and parameters. Furthermore using the concept of a Markov Blanket we can identify all the variables that shield off the class variable from the influence of the remaining network. Therefore Bayesian networks automatically perform feature selection by identifying the (in)dependency relationships with the class variable. RESULTS: We evaluated three methods for integrating clinical and microarray data: decision integration, partial integration and full integration and used them to classify publicly available data on breast cancer patients into a poor and a good prognosis group. The partial integration method is most promising and has an independent test set area under the ROC curve of 0.845. After choosing an operating point the classification performance is better than frequently used indices.

Bayes Theorem↗