Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian networks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Model-based biosignal interpretation.

Two relatively new approaches to model-based biosignal interpretation, qualitative simulation and modelling by causal probabilistic networks, are compared to modelling by differential equations. A major problem in applying a model to an individual patient is the estimation of the parameters. The available observations are unlikely to allow a proper estimation of the parameters, and even if they do, the task appears to have exponential computational complexity if the model is non-linear. Causal probabilistic networks have both differential equation models and qualitative simulation as special cases, and they can provide both Bayesian and maximum-likelihood parameter estimates, in most cases in much less than exponential time. In addition, they can calculate the probabilities required for a decision-theoretical approach to medical decision support. The practical applicability of causal probabilistic networks to real medical problems is illustrated by a model of glucose metabolism which is used to adjust insulin therapy in type I diabetic patients.

Bayes Theorem↗

Mining microarray data to identify transcription factors expressed in naïve resting but not activated T lymphocytes.

Transcriptional repressors controlling the expression of cytokine genes have been implicated in a variety of physiological and pathological phenomena. An unknown repressor that binds to the distal NFAT element of the interleukin-2 (IL-2) gene promoter in naive T-helper lymphocytes has been implicated in autoimmune phenomena and has emerged as a potentially important factor controlling the latency of HIV-1. The aim of this paper was the identification of this repressor. We resorted to public microarray databases looking for DNA-binding proteins that are present in naïve resting T cells but are downregulated when the cells are activated. A Bayesian data mining statistical analysis uncovered 25 candidate factors. Of the 25, NFAT4 and the oncogene ets-2 bind to the common motif AAGGAG found in the HIV-1 LTR and IL-2 probes. Ets-2 binding site contains the three G's that have been shown to be important for binding of the unknown factor; hence, we considered it the likeliest candidate. Electrophoretic mobility shift assays confirmed cross-reactivity between the unknown repressor and anti-ets-2 antibodies, and cotransfection experiments demonstrated the direct involvement of Ets-2 in silencing the IL-2 promoter. Designing experiments for transcription factor analysis using microarrays and Bayesian statistical methodologies provides a novel way toward elucidation of gene control networks.

CD4-Positive T-Lymphocytes↗

Inferring gene regulatory networks with time delays using a genetic algorithm.

Recently a state-space model with time delays for inferring gene regulatory networks was proposed. It was assumed that each regulation between two internal state variables had multiple time delays. This assumption caused underestimation of the model with many current gene expression datasets. In biological reality, one regulatory relationship may have just a single time delay, and not multiple time delays. This study employs Boolean variables to capture the existence of the time-delayed regulatory relationships in gene regulatory networks in terms of the state-space model. As the solution space of time delayed relationships is too large for an exhaustive search, a genetic algorithm (GA) is proposed to determine the optimal Boolean variables (the optimal time-delayed regulatory relationships). Coupled with the proposed GA, Bayesian information criterion (BIC) and probabilistic principle component analysis (PPCA) are employed to infer gene regulatory networks with time delays. Computational experiments are performed on two real gene expression datasets. The results show that the GA is effective at finding time-delayed regulatory relationships. Moreover, the inferred gene regulatory networks with time delays from the datasets improve the prediction accuracy and possess more of the expected properties of a real network, compared to a gene regulatory network without time delays.

Algorithms↗

Stochastic complexities of reduced rank regression in Bayesian estimation.

Reduced rank regression extracts an essential information from examples of input-output pairs. It is understood as a three-layer neural network with linear hidden units. However, reduced rank approximation is a non-regular statistical model which has a degenerate Fisher information matrix. Its generalization error had been left unknown even in statistics. In this paper, we give the exact asymptotic form of its generalization error in Bayesian estimation, based on resolution of learning machine singularities. For this purpose, the maximum pole of the zeta function for the learning theory is calculated. We propose a new method of recursive blowing-ups which yields the complete desingularization of the reduced rank approximation.

Artificial Intelligence↗

Improving Cox survival analysis with a neural-Bayesian approach.

In this article we show that traditional Cox survival analysis can be improved upon when supplemented with sensible priors and analysed within a neural Bayesian framework. We demonstrate that the Bayesian method gives more reliable predictions, in particular for relatively small data sets. The obtained posterior (the probability distribution of network parameters given the data) which in itself is intractable, can be made accessible by several approximations. We review approximations by Hybrid Markov Chain Monte Carlo sampling, a variational method and the Laplace approximation. We argue that although each Bayesian approach circumvents the shortcomings of the original Cox analysis, and therefore yields better predictive results, in practice the use of variational methods or Laplace is preferable. Since Cox survival analysis is infamous for its poor results with (too) many inputs, we use the Bayesian posterior to estimate p-values on the inputs and to formulate an algorithm for backward elimination. We show that after removal of irrelevant inputs Bayesian methods still achieve significantly better results than classical Cox.

Antineoplastic Agents↗

A study of early stopping and model selection applied to the papermaking industry.

This paper addresses the issues of neural network model development and maintenance in the context of a complex task taken from the papermaking industry. In particular, it describes a comparison study of early stopping techniques and model selection, both to optimise neural network models for generalisation performance. The results presented here show that early stopping via use of a Bayesian model evidence measure is a viable way of optimising performance while also making maximum use of all the data. In addition, they show that ten-fold cross-validation performs well as a model selector and as an estimator of prediction accuracy. These results are important in that they show how neural network models may be optimally trained and selected for highly complex industrial tasks where the data are noisy and limited in number.

Algorithms↗

Fine-Scale Landscape Genomics Show Asymmetric Patterns of Gene Flow for the Invasive Mosquito Aedes albopictus.

Mosquito-borne viruses like dengue, Zika, and chikungunya pose increasing health risks in the United States due to the expanding range of Aedes albopictus, a highly invasive mosquito species that now has a global distribution. Aedes albopictus thrive in artificial containers associated with anthropogenic land use, allowing populations to reach high numbers in urban and suburban environments. While the global spread of Ae. albopictus has been well characterized, the effects of heterogeneous urban landscapes on dispersal and gene flow at fine spatial scales remain unclear. This study analyzed the genetic connectivity of Aedes albopictus populations collected in Wake County, North Carolina in 2018. We used single nucleotide polymorphisms (SNP) data from double-digest restriction-enzyme associated DNA sequencing (ddRADseq) and examined genetic connectivity through principal component analysis (PCA) and genetic network analysis. We then evaluated migration and source-sink dynamics using a Bayesian approach for SNP data (BA3-SNP). We found little evidence of genetic clustering or isolated populations of Ae. albopictus in Wake County, suggesting high gene flow between sites. Migration analysis demonstrated asymmetric gene flow from rural to urban regions within Wake County, with greater gene flow occurring between and within urban regions. These findings suggest that the pattern of gene flow of Ae. albopictus populations within local metropolitan areas may involve urban city centers serving as genetic sinks and surrounding suburban and rural regions serving as sources. This study highlights how heterogeneous landscapes shape mosquito population connectivity and migration at fine spatial scales, which is critical for informing vector control and public health intervention strategies.

Aedes albopictus↗

Comparing artificial and convolutional neural networks with traditional models for Genomic prediction in wheat.

With the rapid development of sequencing technology, the application of genomic prediction has become more and more common in breeding schemes of livestocks and crops. Selecting an appropriate statistical model is of central importance to achieve high prediction accuracy. Recently, machine learning models have been expected to upgrade genomic prediction into a new era. However, the perspective still suffers from lack of evidence that machine learning models can generally outperform the traditional ones on empirical data sets. In this study, we compared two machine learning models based on artificial neural network (ANN) and convolutional neural network (CNN) with four traditional models, including genomic best linear unbiased prediction (GBLUP), Bayesian ridge regression (BRR), BayesA and BayesB, using three published data sets for grain yield in wheat. For each model, we considered two variants: modeling and ignoring the genotype-by-environment ([Formula: see text]) interaction. In the comparison, we considered two strategies of cross-validation: predicting genotypes that have not been evaluated in any environment (CV1) and predicting genotypes that have been tested in other environments (CV2). Our results showed that traditional Bayesian models (BayesA, BayesB, and BRR) outperformed GBLUP, ANN and CNN when considering [Formula: see text] interaction. The accuracies of ANN and CNN were higher than traditional models only in CV1 and when [Formula: see text] interaction was ignored. It was also found that the performance of the two machine learning models was significantly affected by the interaction between the CV strategy and the way of treating the [Formula: see text] interaction, while that of the four traditional models was only influenced by whether the [Formula: see text] interaction was considered or not. Thus, machine learning models can be a powerful complementary to the traditional ones and their superiority may depend on the prediction scenario. Among the two machine learning models, we observed that the accuracy of ANN was higher than CNN in most cases, indicating that it is still challenging to adapt complex machine learning models such as CNN to genomic prediction.

ANN↗

An epidemiologic approach to computerized medical diagnosis--AEDMI program.

A program called "An Epidemiological Approach to Computerized Medical Diagnosis" (AEDMI) is presented. Using an interactive questionnaire, physician-patient interviews are conducted and a summary of the relevant clinical data is provided. Standard items, obtained on a multi-centre basis, form a large-scale data base. Simultaneously, the reasoning of clinical experts in each real case is analyzed to obtain a knowledge-rules data base. The methodology of the program combines Bayesian systems, expert systems, and other new lines of research such as neural networks or case-based reasoning. The general concepts of clinical decision making aid systems are reviewed. This publication is aimed at obtaining international cooperation.

Artificial Intelligence↗

Cardiac surgery risk models: a position article.

Differences in medical outcomes may result from disease severity, treatment effectiveness, or chance. Because most outcome studies are observational rather than randomized, risk adjustment is necessary to account for case mix. This has usually been accomplished through the use of standard logistic regression models, although Bayesian models, hierarchical linear models, and machine-learning techniques such as neural networks have also been used. Many factors are essential to insuring the accuracy and usefulness of such models, including selection of an appropriate clinical database, inclusion of critical core variables, precise definitions for predictor variables and endpoints, proper model development, validation, and audit. Risk models may be used to assess the impact of specific predictors on outcome, to aid in patient counseling and treatment selection, to profile provider quality, and to serve as the basis of continuous quality improvement activities.

Bayes Theorem↗

The Kautsky curve is a built-in barcode.

We identify objects from their visually observable morphological features. Automatic methods for identifying living objects are often needed in new technology, and these methods try to utilize shapes. When it comes to identifying plant species automatically, machine vision is difficult to implement because the shapes of different plants overlap and vary greatly because of different viewing angles in field conditions. In the present study we show that chlorophyll a fluorescence, emitted by plant leaves, carries information that can be used for the identification of plant species. Transient changes in fluorescence intensity when a light is turned on were parameterized and then subjected to a variety of pattern recognition procedures. A Self-Organizing Map constructed from the fluorescence signals was found to group the signals according to the phylogenetic origins of the plants. We then used three different methods of pattern recognition, of which the Bayesian Minimum Distance classifier is a parametric technique, whereas the Multilayer Perceptron neural network and k-Nearest Neighbor techniques are nonparametric. Of these techniques, the neural network turned out to be the most powerful one for identifying individual species or groups of species from their fluorescence transients. The excellent recognition accuracy, generally over 95%, allows us to speculate that the method can be further developed into an application in precision agriculture as a means of automatically identifying plant species in the field.

Biophysical Phenomena↗

Phylogenomics and molecular evolution of polyomaviruses.

We provide in this chapter an overview of the basic steps to reconstruct evolutionary relationships through standard phylogeny estimation approaches as well as network approaches for sequences more closely related. We discuss the importance of sequence alignment, selecting models of evolution, and confidence assessment in phylogenetic inference. We also introduce the reader to a variety of software packages used for such studies. Finally, we demonstrate these approaches throughout using a data set of 33 whole genomes of polyomaviruses. A robust phylogeny of these genomes is estimated and phylogenetic relationships among the polyomaviruses determined using Bayesian and maximum likelihood approaches. Furthermore, population samples of SV40 are used to demonstrate the utility of network approaches for closely related sequences. The phylogenetic analysis suggested a close relationship among the BK viruses, JC viruses, and SV40 with a more distant association with mouse polyomavirus, monkey polymavirus (LPV) and then avian polyomavirus (BFDV).

Computational Biology↗

Probabilistic motion estimation based on temporal coherence.

We develop a theory for the temporal integration of visual motion motivated by psychophysical experiments. The theory proposes that input data are temporally grouped and used to predict and estimate the motion flows in the image sequence. This temporal grouping can be considered a generalization of the data association techniques that engineers use to study motion sequences. Our temporal grouping theory is expressed in terms of the Bayesian generalization of standard Kalman filtering. To implement the theory, we derive a parallel network that shares some properties of cortical networks. Computer simulations of this network demonstrate that our theory qualitatively accounts for psychophysical experiments on motion occlusion and motion outliers. In deriving our theory, we assumed spatial factorizability of the probability distributions and made the approximation of updating the marginal distributions of velocity at each point. This allowed us to perform local computations and simplified our implementation. We argue that these approximations are suitable for the stimuli we are considering (for which spatial coherence effects are negligible).

Bayes Theorem↗

Agreement between artificial neural networks and experienced electrocardiographer on electrocardiographic diagnosis of healed myocardial infarction.

OBJECTIVES: The purpose of this study was to compare the diagnoses of healed myocardial infarction made from the 12-lead electrocardiogram (ECG) by artificial neural networks and an experienced electrocardiographer. BACKGROUND: Artificial neural networks have proved of value in pattern recognition tasks. Studies of their utility in ECG interpretation have shown performance exceeding that of conventional ECG interpretation programs. The latter present verbal statements, often with an indication of the likelihood for a certain diagnosis, such as "possible left ventricular hypertrophy." A neural network presents its output as a numeric value between 0 and 1; however, these values can be interpreted as Bayesian probabilities. METHODS: The study was based on 351 healthy volunteers and 1,313 patients with a history of chest pain who had undergone diagnostic cardiac catheterization. A 12-lead ECG was recorded in each subject. An expert electrocardiographer classified the ECGs in five different groups by estimating the probability of anterior myocardial infarction. Artificial neural networks were trained and tested to diagnose anterior myocardial infarction. The network outputs were divided into five groups by using the output values and four thresholds between 0 and 1. RESULTS: The neural networks diagnosed healed anterior myocardial infarctions at high levels of sensitivity and specificity. The network outputs were transformed to verbal statements, and the agreement between these probability estimates and those of an expert electrocardiographer was high. CONCLUSIONS: Artificial neural networks can be of value in automated interpretation of ECGs in the near future.

Electrocardiography↗

Database search post-processing by neural network: Advanced facilities for identification of components in protein mixtures using mass spectrometric peptide mapping.

Database search post-processing by neural network was employed in peptide mapping experiments. The database search was performed using both the known algorithms and score functions, such as Bayesian, MOWSE, Z-score, correlations between calculated and actual peptide length fractional abundance, and, in addition, the probability of protein digest pattern in peptide fingerprint, all embedded in locally developed program. The new signal-processing algorithm based on neural network improves signal-noise separation and is acceptable for automatic protein identification in mixtures. Its power was tested on Helicobacter pylori protein inventory after preceding protein separation by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE). Increase in protein identification success rate was observed, and about 100 proteins were identified with no need of human participation in database search estimation.

Algorithms↗

Algebraic analysis for nonidentifiable learning machines.

This article clarifies the relation between the learning curve and the algebraic geometrical structure of an unidentifiable learning machine such as a multilayer neural network whose true parameter set is an analytic set with singular points. By using a concept in algebraic analysis, we rigorously prove that the Bayesian stochastic complexity or the free energy is asymptotically equal to lambda(1) log n - (m(1) - 1) log log n + constant, where n is the number of training samples and lambda(1) and m(1) are the rational number and the natural number, which are determined as the birational invariant values of the singularities in the parameter space. Also we show an algorithm to calculate lambda(1) and m(1) based on the resolution of singularities in algebraic geometry. In regular statistical models, 2lambda(1) is equal to the number of parameters and m(1) = 1, whereas in nonregular models, such as multilayer networks, 2lambda(1) is not larger than the number of parameters and m(1) > or = 1. Since the increase of the stochastic complexity is equal to the learning curve or the generalization error, the nonidentifiable learning machines are better models than the regular ones if Bayesian ensemble learning is applied.

Algorithms↗

Bayesian fluorescence in situ hybridisation signal classification.

Previous research has indicated the significance of accurate classification of fluorescence in situ hybridisation (FISH) signals for the detection of genetic abnormalities. Based on well-discriminating features and a trainable neural network (NN) classifier, a previous system enabled highly-accurate classification of valid signals and artefacts of two fluorophores. However, since this system employed several features that are considered independent, the naive Bayesian classifier (NBC) is suggested here as an alternative to the NN. The NBC independence assumption permits the decomposition of the high-dimensional likelihood of the model for the data into a product of one-dimensional probability densities. The naive independence assumption together with the Bayesian methodology allow the NBC to predict a posteriori probabilities of class membership using estimated class-conditional densities in a close and simple form. Since the probability densities are the only parameters of the NBC, the misclassification rate of the model is determined exclusively by the quality of density estimation. Densities are evaluated by three methods: single Gaussian estimation (SGE; parametric method), Gaussian mixture model assuming spherical covariance matrices (GMM; semi-parametric method) and kernel density estimation (KDE; non-parametric method). For low-dimensional densities, the GMM generally outperforms the KDE that tends to overfit the training set at the cost of reduced generalisation capability. But, it is the GMM that loses some accuracy when modelling higher-dimensional densities due to the violation of the assumption of spherical covariance matrices when dependent features are added to the set. Compared with these two methods, the SGE and NN provide inferior and superior performance, respectively. However, the NBC avoids the intensive training and optimisation required for the NN, demanding extensive resources and experimentation. Therefore, when supporting these two classifiers, the system enables a trade-off between the NN performance and NBC simplicity of implementation.

Algorithms↗

An MLP training algorithm taking into account known errors on inputs and outputs.

A training algorithm is introduced that takes into account a priori known errors on both inputs and outputs in an MLP network. The new cost function introduced for this case is based on a linear approximation of the network function over the input distribution for a given input pattern. Update formulas, in the form of the gradient of the new cost function, is given for a MLP network, together with expressions for the Hessian matrix. This is later used to calculate error bars in a Bayesian framework. The error bars thus derived are discussed in relation to the more commonly used width of the target posterior predictive distribution. It will also be shown that the taking into account of known input uncertainties in the way suggested in this article will have a strong regularizing effect on the solution.

Algorithms↗