Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian networks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Using fragment chemistry data mining and probabilistic neural networks in screening chemicals for acute toxicity to the fathead minnow.

The paper is illustrating how the general data mining methodology may be adapted to provide solutions to the problem of high throughput virtual screening of organic chemicals for possible acute toxicity to the fathead minnow fish. The present approach involves mining fragment information from chemical structures and is using probabilistic neural networks to model the relationship between structure and toxicity. Probabilistic neural networks implement a special class of multivariate non-linear Bayesian statistical models. The mathematical principles supporting their use for value prediction purposes are clarified and their peculiarities discussed. As part of the research phase of the data mining process, a dataset consisting of 800 structures and associated fathead minnow (Pimephales promelas) 96-h LC50 acute toxicity endpoint information is used for both the purpose of identifying an advantageous combination of fragment descriptors and for training the neural networks. As a result, two powerful models are generated. Model 1 implements the basic PNN with Gaussian kernel (statistical corrections included) while Model 2 implements the PNN with Gaussian kernel and separated variables. External validation is performed using a separate dataset consisting of 86 structures and associated toxicity information. Both learning and generalization capabilities of the two models are investigated and their limitations discussed.

Animals↗

Analyzing cellular biochemistry in terms of molecular networks.

One way to understand cells and circumscribe the function of proteins is through molecular networks. These networks take a variety of forms including webs of protein-protein interactions, regulatory circuits linking transcription factors and targets, and complex pathways of metabolic reactions. We first survey experimental techniques for mapping networks (e.g., the yeast two-hybrid screens). We then turn our attention to computational approaches for predicting networks from individual protein features, such as correlating gene expression levels or analyzing sequence coevolution. All the experimental techniques and individual predictions suffer from noise and systematic biases. These problems can be overcome to some degree through statistical integration of different experimental datasets and predictive features (e.g., within a Bayesian formalism). Next, we discuss approaches for characterizing the topology of networks, such as finding hubs and analyzing subnetworks in terms of common motifs. Finally, we close with perspectives on how network analysis represents a preliminary step toward a systems approach for modeling cells.

Biochemical Phenomena↗

Statistical evaluation of pairwise protein sequence comparison with the Bayesian bootstrap.

MOTIVATION: Protein sequence comparison methods are routinely used to infer the intricate network of evolutionary relationships found within the rapidly growing library of protein sequences, and thereby to predict the structure and function of uncharacterized proteins. In the present study, we detail an improved statistical benchmark of pairwise protein sequence comparison algorithms. We use bootstrap resampling techniques to determine standard statistical errors and to estimate the confidence of our conclusions. We show that the underlying structure within benchmark databases causes Efron's standard, non-parametric bootstrap to be biased. Consequently, the standard bootstrap underpredicts average performance when used in the context of evaluating sequence comparison methods. We have developed, as an alternative, an unbiased statistical evaluation based on the Bayesian bootstrap, a resampling method operationally similar to the standard bootstrap. RESULTS: We apply our analysis to the comparative study of amino acid substitution matrix families and find that using modern matrices results in a small, but statistically significant improvement in remote homology detection compared with the classic PAM and BLOSUM matrices. AVAILABILITY: The sequence sets and code for performing these analyses are available from http://compbio.berkeley.edu/. CONTACT: brenner@compbio.berkeley.edu.

Algorithms↗

Synergistic techniques for better understanding and classifying the environmental structure of landscapes.

The desire to capture natural regions in the landscape has been a goal of geographic and environmental classification and ecological land classification (ELC) for decades. Since the increased adoption of data-centric, multivariate, computational methods, the search for natural regions has become the search for the best classification that optimally trades off classification complexity for class homogeneity. In this study, three techniques are investigated for their ability to find the best classification of the physical environments of the Mt. Lofty Ranges in South Australia: AutoClass-C (a Bayesian classifier), a Kohonen Self-Organising Map neural network, and a k-means classifier with homogeneity analysis. AutoClass-C is specifically designed to find the classification that optimally trades off classification complexity for class homogeneity. However, AutoClass analysis was not found to be assumption-free because it was very sensitive to the user-specified level of relative error of input data. The AutoClass results suggest that there may be no way of finding the best classification without making critical assumptions as to the level of class heterogeneity acceptable in the classification when using continuous environmental data. Therefore, rather than relying on adjusting abstract parameters to arrive at a classification of suitable complexity, it is better to quantify and visualize the data structure and the relationship between classification complexity and class homogeneity. Individually and when integrated, the Self-Organizing Map and k-means classification with homogeneity analysis techniques also used in this study facilitate this and provide information upon which the decision of the scale of classification can be made. It is argued that instead of searching for the elusive classification of natural regions in the landscape, it is much better to understand and visualize the environmental structure of the landscape and to use this knowledge to select the best ELC at the required scale of analysis.

Bayes Theorem↗

Systems biology for cancer.

PURPOSE OF REVIEW: Significant insight can be gained into complex biologic mechanisms of cancer via a combined computational and experimental systems biology approach. This review highlights some of the major systems biology efforts that were applied to cancer in the past year. RECENT FINDINGS: Two main approaches to computational systems biology are discussed: mechanistic dynamical simulations and inferential data mining. Significant developments have occurred in both areas. For example, mechanistic simulations of the EGFR pathway are promoting understanding of cancer, and Bayesian inference approaches allow for the reconstruction of regulatory networks. In addition, the article reports on advancements in experimental systems biology for determining protein-protein interactions and quantifying protein expression to generate the necessary data for computational modeling and inferential data mining. Emerging approaches will further improve the ability to bridge the gap between in vitro systems and in vivo human biology. Technologies paving the way include in vitro models that better reflect in vivo tumors, microfabricated devices of human physiology, and improved animal models. SUMMARY: An important challenge facing the field is how better to translate in vitro discoveries to the clinic. Computational systems biology approaches that use omic data to predict biology along with novel experimental systems that better represent human in vivo biology will prove useful in bridging this gap. Although still early, the potential application of systems biology and the future evolution of the field will significantly affect understanding of cancer disease mechanisms and the ability to devise effective therapeutics.

Computational Biology↗

Machine learning in prognosis of the femoral neck fracture recovery.

We compare the performance of several machine learning algorithms in the problem of prognostics of the femoral neck fracture recovery: the K-nearest neighbours algorithm, the semi-naive Bayesian classifier, backpropagation with weight elimination learning of the multilayered neural networks, the LFC (lookahead feature construction) algorithm, and the Assistant-I and Assistant-R algorithms for top down induction of decision trees using information gain and RELIEFF as search heuristics, respectively. We compare the prognostic accuracy and the explanation ability of different classifiers. Among the different algorithms the semi-naive Bayesian classifier and Assistant-R seem to be the most appropriate. We analyze the combination of decisions of several classifiers for solving prediction problems and show that the combined classifier improves both performance and the explanation ability.

Algorithms↗

A regularization approach to continuous learning with an application to financial derivatives pricing.

We consider the training of neural networks in cases where the nonlinear relationship of interest gradually changes over time. One possibility to deal with this problem is by regularization where a variation penalty is added to the usual mean squared error criterion. To learn the regularized network weights we suggest the Iterative Extended Kalman Filter (IEKF) as a learning rule, which may be derived from a Bayesian perspective on the regularization problem. A primary application of our algorithm is in financial derivatives pricing, where neural networks may be used to model the dependency of the derivatives' price on one or several underlying assets. After giving a brief introduction to the problem of derivatives pricing we present experiments with German stock index options data showing that a regularized neural network trained with the IEKF outperforms several benchmark models and alternative learning procedures. In particular, the performance may be greatly improved using a newly designed neural network architecture that accounts for no-arbitrage pricing restrictions.

Journal Article↗

GenSo-FDSS: a neural-fuzzy decision support system for pediatric ALL cancer subtype identification using gene expression data.

OBJECTIVE: Acute lymphoblastic leukemia (ALL) is the most common malignancy of childhood, representing nearly one third of all pediatric cancers. Currently, the treatment of pediatric ALL is centered on tailoring the intensity of the therapy applied to a patient's risk of relapse, which is linked to the type of leukemia the patient has. Hence, accurate and correct diagnosis of the various leukemia subtypes becomes an important first step in the treatment process. Recently, gene expression profiling using DNA microarrays has been shown to be a viable and accurate diagnostic tool to identify the known prognostically important ALL subtypes. Thus, there is currently a huge interest in developing autonomous classification systems for cancer diagnosis using gene expression data. This is to achieve an unbiased analysis of the data and also partly to handle the large amount of genetic information extracted from the DNA microarrays. METHODOLOGY: Generally, existing medical decision support systems (DSS) for cancer classification and diagnosis are based on traditional statistical methods such as Bayesian decision theory and machine learning models such as neural networks (NN) and support vector machine (SVM). Though high accuracies have been reported for these systems, they fall short on certain critical areas. These included (a) being able to present the extracted knowledge and explain the computed solutions to the users; (b) having a logical deduction process that is similar and intuitive to the human reasoning process; and (c) flexible enough to incorporate new knowledge without running the risk of eroding old but valid information. On the other hand, a neural fuzzy system, which is synthesized to emulate the human ability to learn and reason in the presence of imprecise and incomplete information, has the ability to overcome the above-mentioned shortcomings. However, existing neural fuzzy systems have their own limitations when used in the design and implementation of DSS. Hence, this paper proposed the use of a novel neural fuzzy system: the generic self-organising fuzzy neural network (GenSoFNN) with truth-value restriction (TVR) fuzzy inference, as a fuzzy DSS (denoted as GenSo-FDSS) for the classification of ALL subtypes using gene expression data. RESULTS AND CONCLUSION: The performance of the GenSo-FDSS system is encouraging when benchmarked against those of NN, SVM and the K-nearest neighbor (K-NN) classifier. On average, a classification rate of above 90% has been achieved using the GenSo-FDSS system.

Algorithms↗

Modified Fuzzy ARTMAP Approaches Bayes Optimal Classification Rates: An Empirical Demonstration.

This paper investigates the effectiveness of the Fuzzy ARTMAP (FAM) neural network in classifying statistical data and compares the results with Bayesian decision theory. Binary classification problems are used to assess the performance of FAM operating autonomously and on-line in statistical settings. The results illustrate the limitations of FAM in this context. Novel modifications are, therefore, proposed for the category formation process and the category selection process of FAM, which allow the modified system to minimize the misclassification rates. A number of simulations with randomly generated data sets have been carried out. First, two continuous-valued Gaussian sources are used with various source (mean) separations, prior probabilities, and variances. Then, multi-dimensional discrete patterns are employed to examine the classification ability of modified FAM in both stationary and non-stationary environments. Simulation results consistently demonstrate that modified FAM is able to approach the Bayes optimal classification rates on-line, and thereby justify the rationale behind the modifications. Copyright 1997 Elsevier Science Ltd.

Journal Article↗

Discovery of gene-regulation pathways using local causal search.

This paper reports the methods and results of a computer-based algorithm that takes as input the expression levels of a set of genes as given by DNA microarray data, and then searches for causal pathways that represent how the genes regulate each other. The algorithm uses local heuristic search and a Bayesian scoring metric. We applied the algorithm to induce causal networks from a mixture of observational and experimental gene-expression data on genes involved in galactose metabolism in the yeast Saccharomyces cerevisiae. The observational data consisted of gene-expression levels obtained from unmanipulated inverted exclamation mark degrees wild-type inverted exclamation mark +/- cells. The experimental data were produced by deleting ( inverted exclamation mark degrees knocking out inverted exclamation mark +/-) genes and measuring the expression levels of other genes. We used this data to evaluate several variations of the local search method. In each evaluation, causal relationships were predicted for all 36 pairwise combinations of nine key galactose-related genes. These predictions were then compared to the known causal relationships among these genes.

Algorithms↗

DIAMED: a probabilistic diagnostic aid system on the web.

DIAMED is a system to assist the physicians in the diagnostic process using probabilistic networks as knowledge representation. These networks make it possible to reason on medical data by applying Bayesian methods and to take into account uncertainties of the facts in the resolution of the clinical cases. The proposed model re-uses knowledge contained in an existing knowledge base (ADM). An interface of DIAMED developed on a Web server remotely assists the experts of each medical specialty in updating and validating the knowledge base. Most of the data processing is automated while being based on information preexistent in the ADM base : Constitution of lexicons starting from the existing dictionaries of the ADM system, are then used to work out the requests for selection and update of the knowledge base. One of its assets resides in its pseudo-segmented structure in several layers. The propagation of information is thus limited to only one part of the probabilistic network and calculations are therefore limited.

Artificial Intelligence↗

Generative models for discovering sparse distributed representations.

We describe a hierarchical, generative model that can be viewed as a nonlinear generalization of factor analysis and can be implemented in a neural network. The model uses bottom-up, top-down and lateral connections to perform Bayesian perceptual inference correctly. Once perceptual inference has been performed the connection strengths can be updated using a very simple learning rule that only requires locally available information. We demonstrate that the network learns to extract sparse, distributed, hierarchical representations.

Algorithms↗

Radial Basis Function Networks: Generalization in Over-realizable and Unrealizable Scenarios.

Learning and generalization in a two-layer radial basis function network, with fixed centres of the basis functions, is examined within a stochastic training paradigm. Employing a Bayesian approach, expressions for generalization error are derived under the assumption that the generating mechanism (teacher) for the training data is also a radial basis function network, but one for which the basis function centres and widths need not correspond to those of the student network. The effects of regularization, via a weight decay term, are examined. The cases in which the student has greater representational power than the teacher (over-realizable), and in which the teacher has greater power than the student (unrealizable) are studied. Dependence on knowing the centres of the teacher is eliminated by introducing a single degree-of-confidence parameter. Finally, simulations are performed which validate the analytic results. Copyright 1996 Elsevier Science Ltd.

Journal Article↗

Modeling genetic networks from clonal analysis.

In this report a systematic approach is used to determine the approximate genetic network and robust dependencies underlying differentiation. The data considered is in the form of a binary matrix and represent the expression of the nine genes across the 99 colonies. The report is divided into two parts: the first part identifies significant pair-wise dependencies from the given binary matrix using linear correlation and mutual information. A new method is proposed to determine statistically significant dependencies estimated using the mutual information measure. In the second, a Bayesian approach is used to obtain an approximate description (equivalence class) of network structures. The robustness of linear correlation, mutual information and the equivalence class of networks is investigated with perturbation and decreasing colony number. Perturbation of the data was achieved by generating bootstrap realizations. The results are refined with biological knowledge. It was found that certain dependencies in the network are immune to perturbation and decreasing colony number and may represent robust features, inherent in the differentiation program of osteoblast progenitor cells. The methods to be discussed are generic in nature and not restricted to the experimental paradigm addressed in this study.

Animals↗

Bringing metabolic networks to life: integration of kinetic, metabolic, and proteomic data.

BACKGROUND: Translating a known metabolic network into a dynamic model requires reasonable guesses of all enzyme parameters. In Bayesian parameter estimation, model parameters are described by a posterior probability distribution, which scores the potential parameter sets, showing how well each of them agrees with the data and with the prior assumptions made. RESULTS: We compute posterior distributions of kinetic parameters within a Bayesian framework, based on integration of kinetic, thermodynamic, metabolic, and proteomic data. The structure of the metabolic system (i.e., stoichiometries and enzyme regulation) needs to be known, and the reactions are modelled by convenience kinetics with thermodynamically independent parameters. The parameter posterior is computed in two separate steps: a first posterior summarises the available data on enzyme kinetic parameters; an improved second posterior is obtained by integrating metabolic fluxes, concentrations, and enzyme concentrations for one or more steady states. The data can be heterogeneous, incomplete, and uncertain, and the posterior is approximated by a multivariate log-normal distribution. We apply the method to a model of the threonine synthesis pathway: the integration of metabolic data has little effect on the marginal posterior distributions of individual model parameters. Nevertheless, it leads to strong correlations between the parameters in the joint posterior distribution, which greatly improve the model predictions by the following Monte-Carlo simulations. CONCLUSION: We present a standardised method to translate metabolic networks into dynamic models. To determine the model parameters, evidence from various experimental data is combined and weighted using Bayesian parameter estimation. The resulting posterior parameter distribution describes a statistical ensemble of parameter sets; the parameter variances and correlations can account for missing knowledge, measurement uncertainties, or biological variability. The posterior distribution can be used to sample model instances and to obtain probabilistic statements about the model's dynamic behaviour.

Bayes Theorem↗

Supervised classification for gene network reconstruction.

One of the central problems of functional genomics is revealing gene expression networks - the relationships between genes that reflect observations of how the expression level of each gene affects those of others. Microarray data are currently a major source of information about the interplay of biochemical network participants in living cells. Various mathematical techniques, such as differential equations, Bayesian and Boolean models and several statistical methods, have been applied to expression data in attempts to extract the underlying knowledge. Unsupervised clustering methods are often considered as the necessary first step in visualization and analysis of the expression data. As for supervised classification, the problem mainly addressed so far has been how to find discriminative genes separating various samples or experimental conditions. Numerous methods have been applied to identify genes that help to predict treatment outcome or to confirm a diagnosis, as well as to identify primary elements of gene regulatory circuits. However, less attention has been devoted to using supervised learning to uncover relationships between genes and/or their products. To start filling this gap a machine-learning approach for gene networks reconstruction is described here. This approach is based on building classifiers--functions, which determine the state of a gene's transcription machinery through expression levels of other genes. The method can be applied to various cases where relationships between gene expression levels could be expected.

Genes↗

Integration of form and motion within a generative model of visual cortex.

One of the challenges faced by the visual system is integrating cues within and across processing streams for inferring scene properties and structure. This is particularly apparent in the inference of object motion, where psychophysical experiments have shown that integration of motion signals, distributed across space, must also be integrated with form cues. This has led several to conclude that there exist mechanisms which enable form cues to 'veto' or completely suppress ambiguous motion signals. We describe a probabilistic approach which uses a generative network model for integrating form and motion cues using the machinery of belief propagation and Bayesian inference. We show, using computer simulations, that motion integration can be mediated via a local, probabilistic representation of contour ownership, which we have previously termed 'direction of figure'. The uncertainty of this inferred form cue is used to modulate the covariance matrix of network nodes representing local motion estimates in the motion stream. We show with results for two sets of stimuli that the model does not completely suppress ambiguous cues, but instead integrates them in a way that is a function of their underlying uncertainty. The result is that the model can account for the continuum of bias seen for motion coherence and perceived object motion in psychophysical experiments.

Bayes Theorem↗

A hybrid neural and statistical classifier system for histopathologic grading of prostatic lesions.

Neural network and statistical classification methods were applied to derive an objective grading for moderately and poorly differentiated lesions of the prostate, based on characteristics of the nuclear placement patterns. A partly trained multilayer neural network was used as a feature extractor. A hybrid classifier system using a quadratic Bayesian classifier applied to these features allowed grade assignment consensus with visual diagnosis in 96% of fields from a training set of 500 fields and in 77% of 130 fields of a test set.

Humans↗