Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian networks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Modeling of farnesyltransferase inhibition by some thiol and non-thiol peptidomimetic inhibitors using genetic neural networks and RDF approaches.

Inhibition of farnesyltransferase (FT) enzyme by a set of 78 thiol and non-thiol peptidomimetic inhibitors was successfully modeled by a genetic neural network (GNN) approach, using radial distribution function descriptors. A linear model was unable to successfully fit the whole data set; however, the optimum Bayesian regularized neural network model described about 87% inhibitory activity variance with a relevant predictive power measured by q2 values of leave-one-out and leave-group-out cross-validations of about 0.7. According to their activity levels, thiol and non-thiol inhibitors were well-distributed in a topological map, built with the inputs of the optimum non-linear predictor. Furthermore, descriptors in the GNN model suggested the occurrence of a strong dependence of FT inhibition on the molecular shape and size rather than on electronegativity or polarizability characteristics of the studied compounds.

Enzyme Inhibitors↗

A data mining approach for signal detection and analysis.

The WHO database contains over 2.5 million case reports, analysis of this data set is performed with the intention of signal detection. This paper presents an overview of the quantitative method used to highlight dependencies in this data set. The method Bayesian confidence propagation neural network (BCPNN) is used to highlight dependencies in the data set. The method uses Bayesian statistics implemented in a neural network architecture to analyse all reported drug adverse reaction combinations. This method is now in routine use for drug adverse reaction signal detection. Also this approach has been extended to highlight drug group effects and look for higher order dependencies in the WHO data. Quantitatively unexpectedly strong relationships in the data are highlighted relative to general reporting of suspected adverse effects; these associations are then clinically assessed.

Adverse Drug Reaction Reporting Systems↗

Reconstructing genetic networks from time ordered gene expression data using Bayesian method with global search algorithm.

Different genes of an organism are expressed to different levels at different times during the life cycle and in response to various environmental stresses. Elucidating the network of gene-gene interactions responsible for the expression helps understand living processes. Microarray technology allows concurrent genomic scale measurement of an organism's mRNA levels. We describe a power-law formalism to model the combinatorial effect of regulators on gene transcription. The dynamic model allows delayed transcription. We employ a principled network reconstruction approach that accounts for the high noise and low replicate characteristics of present day microarray data. An important feature of our approach is that the detail of the reconstructed network is limited to the noise level of the data. We apply the methodology to a microarray dataset of yeast cells grown in glucose and experiencing a diauxic transition upon glucose depletion. The reconstructed transcriptional regulations of yeast glycolytic genes are consistent with published findings.

Algorithms↗

Bayesian error analysis model for reconstructing transcriptional regulatory networks.

Transcription regulation is a fundamental biological process, and extensive efforts have been made to dissect its mechanisms through direct biological experiments and regulation modeling based on physical-chemical principles and mathematical formulations. Despite these efforts, transcription regulation is yet not well understood because of its complexity and limitations in biological experiments. Recent advances in high throughput technologies have provided substantial amounts and diverse types of genomic data that reveal valuable information on transcription regulation, including DNA sequence data, protein-DNA binding data, microarray gene expression data, and others. In this article, we propose a Bayesian error analysis model to integrate protein-DNA binding data and gene expression data to reconstruct transcriptional regulatory networks. There are two unique aspects to this proposed model. First, transcription is modeled as a set of biochemical reactions, and a linear system model with clear biological interpretation is developed. Second, measurement errors in both protein-DNA binding data and gene expression data are explicitly considered in a Bayesian hierarchical model framework. Model parameters are inferred through Markov chain Monte Carlo. The usefulness of this approach is demonstrated through its application to infer transcriptional regulatory networks in the yeast cell cycle.

Algorithms↗

A comparison of three techniques for rapid model development: an application in patient risk-stratification.

Accurately risk-stratifying patients is a key component of health care outcomes assessment. And, many health care organizations increasingly are relying upon automated means for assistance in making patient risk-stratification decisions. Unfortunately, the process of outcome model development, as it is currently practiced, is both time consuming and difficult. We investigated the relative abilities of three modeling techniques (logistic regression, artificial neural network (ANN), and Bayesian) to rapidly develop models for risk-stratifying patients. Our results demonstrated that all three modeling techniques perform equally well in certain situations. However, the Bayesian model with conditional independence had the best overall performance. Unfortunately, none of the models were able to achieve the degree of accuracy which would be required in a medical setting.

APACHE↗

Medical expert systems based on causal probabilistic networks.

Causal probabilistic networks (CPNs) offer new methods by which you can build medical expert systems that can handle all types of medical reasoning within a uniform conceptual framework. Based on the experience from a commercially available system and a couple of large prototype systems, it appears that CPNs are now an attractive alternative to other methods. A CPN is an intensional model of a domain, and it is therefore conceptually much closer to qualitative reasoning systems and to simulation systems than to rule-based or logic-based systems. Recent progress in Bayesian inference in networks has yielded computationally efficient methods. The inference method used follows the fundamental axioms of probability theory, and gives a sound framework for causal and diagnostic (deductive and abductive) reasoning under uncertainty. Experience with the prototypes indicates that it may be possible to use decision theory as a rational approach to test planning and therapy planning. The way in which knowledge is acquired and represented in CPNs makes it easy to express 'deep knowledge' for example in the form of physiological models, and the facilities for learning make it possible to make a smooth transition from expert opinion to statistics based on empirical data.

Artificial Intelligence↗

A Bayesian regression approach to the inference of regulatory networks from gene expression data.

MOTIVATION: There is currently much interest in reverse-engineering regulatory relationships between genes from microarray expression data. We propose a new algorithmic method for inferring such interactions between genes using data from gene knockout experiments. The algorithm we use is the Sparse Bayesian regression algorithm of Tipping and Faul. This method is highly suited to this problem as it does not require the data to be discretized, overcomes the need for an explicit topology search and, most importantly, requires no heuristic thresholding of the discovered connections. RESULTS: Using simulated expression data, we are able to show that this algorithm outperforms a recently published correlation-based approach. Crucially, it does this without the need to set any ad hoc threshold on possible connections.

Algorithms↗

A Bayesian connectivity-based approach to constructing probabilistic gene regulatory networks.

MOTIVATION: We have hypothesized that the construction of transcriptional regulatory networks using a method that optimizes connectivity would lead to regulation consistent with biological expectations. A key expectation is that the hypothetical networks should produce a few, very strong attractors, highly similar to the original observations, mimicking biological state stability and determinism. Another central expectation is that, since it is expected that the biological control is distributed and mutually reinforcing, interpretation of the observations should lead to a very small number of connection schemes. RESULTS: We propose a fully Bayesian approach to constructing probabilistic gene regulatory networks (PGRNs) that emphasizes network topology. The method computes the possible parent sets of each gene, the corresponding predictors and the associated probabilities based on a nonlinear perceptron model, using a reversible jump Markov chain Monte Carlo (MCMC) technique, and an MCMC method is employed to search the network configurations to find those with the highest Bayesian scores to construct the PGRN. The Bayesian method has been used to construct a PGRN based on the observed behavior of a set of genes whose expression patterns vary across a set of melanoma samples exhibiting two very different phenotypes with respect to cell motility and invasiveness. Key biological features have been faithfully reflected in the model. Its steady-state distribution contains attractors that are either identical or very similar to the states observed in the data, and many of the attractors are singletons, which mimics the biological propensity to stably occupy a given state. Most interestingly, the connectivity rules for the most optimal generated networks constituting the PGRN are remarkably similar, as would be expected for a network operating on a distributed basis, with strong interactions between the components.

Algorithms↗

Linear and nonlinear QSAR study of N-hydroxy-2-[(phenylsulfonyl)amino]acetamide derivatives as matrix metalloproteinase inhibitors.

The inhibitory activity (IC50) toward matrix metalloproteinases (MMP-1, MMP-2, MMP-3, MMP-9, and MMP-13) of N-hydroxy-2-[(phenylsulfonyl)amino]acetamide derivatives (HPSAAs) has been successfully modeled using 2D autocorrelation descriptors. The relevant molecular descriptors were selected by linear and nonlinear genetic algorithm (GA) feature selection using multiple linear regression (MLR) and Bayesian-regularized neural network (BRANN) approaches, respectively. The quality of the models was evaluated by means of cross-validation experiments and the best results correspond to nonlinear ones (Q2>0.7 for all models). Despite the high correlation between the studied compound IC50 values, the 2D autocorrelation space brings different descriptors for each MMP inhibition. On the basis of these results, these models contain useful molecular information about the ligand specificity for MMP S'1, S1, and S'2 pockets.

Acetamides↗

Modeling K(m) values using electrotopological state: substrates for cytochrome P450 3A4-mediated metabolism.

In order to determine K(m) values of substrates for CYP3A4-mediated metabolism, an in silico model has been developed in the present work. Using electrotopological state (E-state) indices, together with Bayesian-regularized neural network (BRNN), we have described an in silico method to model log(1/K(m)) values of various substrates. The relative importance of the E-state indices is analyzed by principal component analysis. By using an additional external test set, which is independent of the training set, the robustness and predictivity of the model are also validated.

Computational Biology↗

Machine learning for medical diagnosis: history, state of the art and perspective.

The paper provides an overview of the development of intelligent data analysis in medicine from a machine learning perspective: a historical view, a state-of-the-art view, and a view on some future trends in this subfield of applied artificial intelligence. The paper is not intended to provide a comprehensive overview but rather describes some subareas and directions which from my personal point of view seem to be important for applying machine learning in medical diagnosis. In the historical overview, I emphasize the naive Bayesian classifier, neural networks and decision trees. I present a comparison of some state-of-the-art systems, representatives from each branch of machine learning, when applied to several medical diagnostic tasks. The future trends are illustrated by two case studies. The first describes a recently developed method for dealing with reliability of decisions of classifiers, which seems to be promising for intelligent data analysis in medicine. The second describes an approach to using machine learning in order to verify some unexplained phenomena from complementary medicine, which is not (yet) approved by the orthodox medical community but could in the future play an important role in overall medical diagnosis and treatment.

Artificial Intelligence↗

Designing libraries with CNS activity.

Library design is an important and difficult task. In this paper we describe one possible solution to designing a CNS-active library. CNS-actives and -inactives were selected from the CMC and the MDDR databases based on whether they were described as having some kind of CNS activity in the databases. This classification scheme results in over 15 000 actives and over 50 000 inactives. Each molecule is described by 7 1D descriptors (molecular weight, number of donors, number of acceptors, etc.) and 166 2D descriptors (presence/absence of functional groups such as NH(2)). A neural network trained using Bayesian methods can correctly predict about 75% of the actives and 65% of the inactives using the 7 1D descriptors. The performance improves to a prediction accuracy on the active set of 83% and 79% on the inactives on adding the 2D descriptors. On a database with 275 compounds where the CNS activity is known (from the literature) for each compound, we achieve 92% and 71% accuracy on the actives and inactives, respectively. The models we construct can therefore be used as a "filter" to examine any set of proposed molecules in a chemical library. As an example of the utility of our method, we describe the generation of a small library of potentially CNS-active molecules that would be amenable to combinatorial chemistry. This was done by building and analyzing a large database of a million compounds constructed from frameworks and side chains frequently found in drug molecules.

Animals↗

Development of quantitative structure-activity relationships for the toxicity of aromatic compounds to Tetrahymena pyriformis: comparative assessment of the methodologies.

The purpose of this study was to develop quantitative structure-activity relationships (QSARs) for the toxicity of 268 aromatic compounds in the Tetrahymena pyriformis growth inhibition assay. The QSARs were developed using the response-surface (or two-parameter) approach, which was also modified using linear free-energy parameters to account for outliers. Subsequently, the data set was analyzed using partial least-squares (PLS). The results of the modeling using different methodologies were compared to the use of a Bayesian regularized neural network (BRANN) trained on the same data. Both response surface approaches, and PLS explained between 75 and 80% of the variance in the data; BRANN gave a higher statistical fit. In terms of the transparency of the approaches, the response surface clearly provides the simplest and easiest to use QSAR, it is readily interpreted in terms of mechanism of toxic action. PLS and BRANN are respectively less transparent. The use of atomistic and fragment-based indexes as descriptors in QSARs is assessed also, these are found not to be as useful as whole molecule parameters for the prediction of toxicity for molecules outside of the training set. The relative merits of the different approaches to the development of QSARs are described.

Animals↗

Extension of a local backbone description using a structural alphabet: a new approach to the sequence-structure relationship.

Protein Blocks (PBs) comprise a structural alphabet of 16 protein fragments, each 5 Calpha long. They make it possible to approximate and correctly predict local protein three-dimensional (3D) structures. We have selected the 72 most frequent sequences of five PBs, which we call Structural Words (SWs). Analysis of four different protein data banks shows that SWs cover 92% of the amino acids in them and provide a good structural approximation for residues (i.e., sequences) 9 Calpha long. We present most of them in a simple network that describes 90% of the overall residues and, interestingly, includes more than 80% of the amino acids present in coils. Analysis of the network shows the specificity and quality of the 3D descriptions as well as a new type of relation between local folds and amino acid distribution. The results show that the 3D structure of these protein data banks can be easily described by a combination of subgraphs included in the network. Finally, a Bayesian probabilistic approach improved the prediction rate by 4%.

Amino Acid Sequence↗

Estimating three-class ideal observer decision variables for computerized detection and classification of mammographic mass lesions.

We are using Bayesian artificial neural networks (BANNs) to classify mammographic masses in schemes for computer-aided diagnosis, and we are extending this methodology to a three-class classification task. We investigated whether a BANN can estimate ideal observer decision variables to distinguish malignant, benign, and false-positive computer detections. Five features were calculated for 63 malignant and 29 benign computer-detected mass lesions, and for 1049 false-positive computer detections, in 440 mammograms randomly divided into a training and testing set. A BANN was trained on the training set features and applied to the testing set features. We then used a known relation between three-class ideal observer decision variables and that used by a two-class ideal observer when two of three classes are grouped into one class, giving one decision variable for distinguishing malignant from nonmalignant detections, and a second for distinguishing true-positive from false-positive computer detections. For comparison, we grouped the training data into two classes in the same two ways and trained two-class BANNs for these two tasks. The three-class BANN decision variables were essentially identical in performance to the specifically trained two-class BANNs, with the average difference in area under the ROC curves being less than 0.0035 and no differences in area being statistically significant. Thus, the BANN outputs obey the same theoretical relationship as do the three-class and two-class ideal observer decision variables, which is consistent with the claim that the three-class BANN output can provide good estimates of the decision variables used by a three-class ideal observer.

Breast Neoplasms↗

Radial gradient-based segmentation of mammographic microcalcifications: observer evaluation and effect on CAD performance.

Precise segmentation of microcalcifications is essential in the development of accurate mammographic computer-aided diagnosis (CAD) schemes. We have designed a radial gradient-based segmentation method for microcalcifications, and compared it to both the region-growing segmentation method currently used in our CAD scheme and to the watershed segmentation method. Two observer studies were conducted to subjectively evaluate the proposed segmentation method. The first study (A) required observers to rate the segmentation accuracy on a 100-point scale. The second observer evaluation (B) was a preference study in which observers selected their preferred method from three displayed segmentation methods. In study A, the observers gave an average accuracy rating of 88 for the radial gradient-based and 50 for the region-growing segmentation method. In study B, the two observers selected the proposed method 56% and 62% of the time. We also investigated the effect of the proposed segmentation method on the performance of computerized classification scheme in differentiating malignant from benign clustered microcalcifications. The performances of the classification scheme using a linear discriminant analysis (LDA) or a Bayesian artificial neural network classifier both showed statistically significant improvements when using the proposed segmentation method. The areas under the receiver-operating characteristic curves for case-based performance when using the LDA classifier were 0.86 with the proposed segmentation method, 0.80 with the region-growing method, and 0.83 with the watershed method.

Algorithms↗

Decision-support and intelligent tutoring systems in medical education.

One of the challenges in medical education is to teach the decision-making process. This learning process varies according to the experience of the student and can be supported by various tools. In this paper we present several approaches that can strengthen this mechanism, from decision-support tools, such as scoring systems, Bayesian models, neural networks, to cognitive models that can reproduce how the students progressively build their knowledge into memory and foster pedagogic methods.

Artificial Intelligence↗

Identifying crash propensity using specific traffic speed conditions.

INTRODUCTION: In spite of recent advances in traffic surveillance technology and ever-growing concern over traffic safety, there have been very few research efforts establishing links between real-time traffic flow parameters and crash occurrence. This study aims at identifying patterns in the freeway loop detector data that potentially precede traffic crashes. METHOD: The proposed solution essentially involves classification of traffic speed patterns emerging from the loop detector data. Historical crash and loop detector data from the Interstate-4 corridor in the Orlando metropolitan area were used for this study. Traffic speed data from sensors embedded in the pavement (i.e., loop detector stations) to measure characteristics of the traffic flow were collected for both crash and non-crash conditions. Bayesian classifier based methodology, probabilistic neural network (PNN), was then used to classify these data as belonging to either crashes or non-crashes. PNN is a neural network implementation of well-known Bayesian-Parzen classifier. With its superb mathematical credentials, the PNN trains much faster than multilayer feed forward networks. The inputs to final classification model, selected from various candidate models, were logarithms of the coefficient of variation in speed obtained from three stations, namely, station of the crash (i.e., station nearest to the crash location) and two stations immediately preceding it in the upstream direction (measured in 5 minute time slices of 10-15 minutes prior to the crash time). RESULTS: The results showed that at least 70% of the crashes on the evaluation dataset could be identified using the classifiers developed in this paper.

Accidents, Traffic↗