Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian networks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Identification of genes associated with ovarian cancer metastasis using microarray expression analysis.

Although the transition from early- to advanced-stage ovarian cancer is a critical determinant of survival, little is known about the molecular underpinnings of ovarian metastasis. We hypothesize that microarray analysis of global gene expression patterns in primary ovarian cancer and metastatic omental implants can identify genes that underlie the metastatic process in epithelial ovarian cancer. We utilized Affymetrix U95Av2 microarrays to characterize the molecular alterations that underlie omental metastasis from 47 epithelial ovarian cancer samples collected from multiple sites in 20 patients undergoing primary surgical cytoreduction for advanced-stage (IIIC/IV) serous ovarian cancer. Fifty-six genes demonstrated differential expression between ovarian and omental samples (P < 0.01), and twenty of these 56 differentially expressed genes have previously been implicated in metastasis, cell motility, or cytoskeletal function. Ten of the 56 genes are involved in p53 gene pathways. A Bayesian statistical tree analysis was used to identify a 27-gene expression pattern that could accurately predict the site of tumor (ovary versus omentum). This predictive model was evaluated using an external data set. Nine of the 27 predictive genes have previously been shown to be involved in oncogenesis and/or metastasis, and 10/27 genes have been implicated in p53 pathways. Microarray findings were validated by real-time quantitative PCR. We conclude that gene expression patterns that distinguish omental metastasis from primary epithelial ovarian cancer can be identified and that many of the genes have functions that are biologically consistent with a role in oncogenesis, metastasis, and p53 gene networks.

Bayes Theorem↗

Sequential Bayesian decoding with a population of neurons.

Population coding is a simplified model of distributed information processing in the brain. This study investigates the performance and implementation of a sequential Bayesian decoding (SBD) paradigm in the framework of population coding. In the first step of decoding, when no prior knowledge is available, maximum likelihood inference is used; the result forms the prior knowledge of stimulus for the second step of decoding. Estimates are propagated sequentially to apply maximum a posteriori (MAP) decoding in which prior knowledge for any step is taken from estimates from the previous step. Not only do we analyze the performance of SBD, obtaining the optimal form of prior knowledge that achieves the best estimation result, but we also investigate its possible biological realization, in the sense that all operations are performed by the dynamics of a recurrent network. In order to achieve MAP, a crucial point is to identify a mechanism that propagates prior knowledge. We find that this could be achieved by short-term adaptation of network weights according to the Hebbian learning rule. Simulation results on both constant and time-varying stimulus support the analysis.

Bayes Theorem↗

Identification of regulatory targets of tissue-specific transcription factors: application to retina-specific gene regulation.

Identification of tissue-specific gene regulatory networks can yield insights into the molecular basis of a tissue's development, function and pathology. Here, we present a computational approach designed to identify potential regulatory target genes of photoreceptor cell-specific transcription factors (TFs). The approach is based on the hypothesis that genes related to the retina in terms of expression, disease and/or function are more likely to be the targets of retina-specific TFs than other genes. A list of genes that are preferentially expressed in retina was obtained by integrating expressed sequence tag, SAGE and microarray datasets. The regulatory targets of retina-specific TFs are enriched in this set of retina-related genes. A Bayesian approach was employed to integrate information about binding site location relative to a gene's transcription start site. Our method was applied to three retina-specific TFs, CRX, NRL and NR2E3, and a number of potential targets were predicted. To experimentally assess the validity of the bioinformatic predictions, mobility shift, transient transfection and chromatin immunoprecipitation assays were performed with five predicted CRX targets, and the results were suggestive of CRX regulation in 5/5, 3/5 and 4/5 cases, respectively. Together, these experiments strongly suggest that RP1, GUCY2D, ABCA4 are novel targets of CRX.

Animals↗

Inference and computation with population codes.

In the vertebrate nervous system, sensory stimuli are typically encoded through the concerted activity of large populations of neurons. Classically, these patterns of activity have been treated as encoding the value of the stimulus (e.g., the orientation of a contour), and computation has been formalized in terms of function approximation. More recently, there have been several suggestions that neural computation is akin to a Bayesian inference process, with population activity patterns representing uncertainty about stimuli in the form of probability distributions (e.g., the probability density function over the orientation of a contour). This paper reviews both approaches, with a particular emphasis on the latter, which we see as a very promising framework for future modeling and experimental work.

Animals↗

Incorporating prior information via shrinkage: a combined analysis of genome-wide location data and gene expression data.

Transcriptional control is a critical step in regulation of gene expression. Understanding such a control on a genomic level involves deciphering the mechanisms and structures of regulatory programmes and networks. A difficulty arises due to the weak signal and high noise in various sources of data while most current approaches are limited to analysis of a single source of data. A natural alternative is to improve statistical efficiency and power by a combined analysis of multiple sources of data. Here we propose a shrinkage method to combine genome-wide location data and gene expression data to detect the binding sites or target genes of a transcription factor. Specifically, a prior 'non-target' gene list is generated by analysing the expression data, and then this information is incorporated into the subsequent binding data analysis via a shrinkage method. There is a Bayesian justification for this shrinkage method. Both simulated and real data were used to evaluate the proposed method and compare it with analysing binding data alone. In simulation studies, the proposed method gives higher sensitivity and lower false discovery rate (FDR) in detecting the target genes. In real data example, the proposed method can reduce the estimated FDR and increase the power to detect the previously known target genes of a broad transcription regulator, leucine responsive regulatory protein (Lrp) in Escherichia coli. This method can also be used to incorporate other information, such as gene ontology (GO), to microarray data analysis to detect differentially expressed genes.

Bayes Theorem↗

Variational mixture of Bayesian independent component analyzers.

There has been growing interest in subspace data modeling over the past few years. Methods such as principal component analysis, factor analysis, and independent component analysis have gained in popularity and have found many applications in image modeling, signal processing, and data compression, to name just a few. As applications and computing power grow, more and more sophisticated analyses and meaningful representations are sought. Mixture modeling methods have been proposed for principal and factor analyzers that exploit local gaussian features in the subspace manifolds. Meaningful representations may be lost, however, if these local features are nongaussian or discontinuous. In this article, we propose extending the gaussian analyzers mixture model to an independent component analyzers mixture model. We employ recent developments in variational Bayesian inference and structure determination to construct a novel approach for modeling nongaussian, discontinuous manifolds. We automatically determine the local dimensionality of each manifold and use variational inference to calculate the optimum number of ICA components needed in our mixture model. We demonstrate our framework on complex synthetic data and illustrate its application to real data by decomposing functional magnetic resonance images into meaningful-and medically useful-features.

Bayes Theorem↗

Variational studies and replica symmetry breaking in the generalization problem of the binary perceptron

We analyze the average performance of a general class of learning algorithms for the nondeterministic polynomial time complete problem of rule extraction by a binary perceptron. The examples are generated by a rule implemented by a teacher network of similar architecture. A variational approach is used in trying to identify the potential energy that leads to the largest generalization in the thermodynamic limit. We restrict our search to algorithms that always satisfy the binary constraints. A replica symmetric ansatz leads to a learning algorithm which presents a phase transition in violation of an information theoretical bound. Stability analysis shows that this is due to a failure of the replica symmetric ansatz and the first step of replica symmetry breaking (RSB) is studied. The variational method does not determine a unique potential but it allows construction of a class with a unique minimum within each first order valley. Members of this class improve on the performance of Gibbs algorithm but fail to reach the Bayesian limit in the low generalization phase. They even fail to reach the performance of the best binary, an optimal clipping of the barycenter of version space. We find a trade-off between a good low performance and early onset of perfect generalization. Although the RSB may be locally stable we discuss the possibility that it fails to be the correct saddle point globally.

Journal Article↗

An electronic tongue using potentiometric all-solid-state PVC-membrane sensors for the simultaneous quantification of ammonium and potassium ions in water.

The simultaneous determination of NH(4)(+) and K(+) in solution has been attempted using a potentiometric sensor array and multivariate calibration. The sensors used are rather non-specific and of all-solid-state type, employing polymeric (PVC) membranes. The subsequent data processing is based on the use of a multilayer artificial neural network (ANN). This approach is given the name "electronic tongue" because it mimics the sense of taste in animals. The sensors incorporate, as recognition elements, neutral carriers belonging to the family of the ionophoric antibiotics. In this work the ANN type is optimized by studying its topology, the training algorithm, and the transfer functions. Also, different pretreatments of the starting data are evaluated. The chosen ANN is formed by 8 input neurons, 20 neurons in the hidden layer and 2 neurons in the output layer. The transfer function selected for the hidden layer was sigmoidal and linear for the output layer. It is also recommended to scale the starting data before training. A correct fit for the test data set is obtained when it is trained with the Bayesian regularization algorithm. The viability for the determination of ammonium and potassium ions in synthetic samples was evaluated; cumulative prediction errors of approximately 1% (relative values) were obtained. These results were comparable with those obtained with a generalized regression ANN as a reference algorithm. In a final application, results close to the expected values were obtained for the two considered ions, with concentrations between 0 and 40 mmol L(-1).

Journal Article↗

Bridging structural biology and genomics: assessing protein interaction data with known complexes.

Currently, there is a major effort to map protein-protein interactions on a genome-wide scale. The utility of the resulting interaction networks will depend on the reliability of the experimental methods and the coverage of the approaches. Known macromolecular complexes provide a defined and objective set of protein interactions with which to compare biochemical and genetic data for validation. Here, we show that a significant fraction of the protein-protein interactions in genome-wide datasets, as well as many of the individual interactions reported in the literature, are inconsistent with the known 3D structures of three recent complexes (RNA polymerase II, Arp2/3 and the proteasome). Furthermore, comparison among genome-wide datasets, and between them and a larger (but less well resolved) group of 174 complexes, also shows marked inconsistencies. Finally, individual interaction datasets, being inherently noisy, are best used when integrated together, and we show how simple Bayesian approaches can combine them, significantly decreasing error rate.

Actin-Related Protein 3↗

Evolution of the cerebellum as a neuronal machine for Bayesian state estimation.

The cerebellum evolved in association with the electric sense and vestibular sense of the earliest vertebrates. Accurate information provided by these sensory systems would have been essential for precise control of orienting behavior in predation. A simple model shows that individual spikes in electrosensory primary afferent neurons can be interpreted as measurements of prey location. Using this result, I construct a computational neural model in which the spatial distribution of spikes in a secondary electrosensory map forms a Monte Carlo approximation to the Bayesian posterior distribution of prey locations given the sense data. The neural circuit that emerges naturally to perform this task resembles the cerebellar-like hindbrain electrosensory filtering circuitry of sharks and other electrosensory vertebrates. The optimal filtering mechanism can be extended to handle dynamical targets observed from a dynamical platform; that is, to construct an optimal dynamical state estimator using spiking neurons. This may provide a generic model of cerebellar computation. Vertebrate motion-sensing neurons have specific fractional-order dynamical characteristics that allow Bayesian state estimators to be implemented elegantly and efficiently, using simple operations with asynchronous pulses, i.e. spikes. The computational neural models described in this paper represent a novel kind of particle filter, using spikes as particles. The models are specific and make testable predictions about computational mechanisms in cerebellar circuitry, while providing a plausible explanation of cerebellar contributions to aspects of motor control, perception and cognition.

Adaptation, Physiological↗

Stochastic models of neuronal dynamics.

Cortical activity is the product of interactions among neuronal populations. Macroscopic electrophysiological phenomena are generated by these interactions. In principle, the mechanisms of these interactions afford constraints on biologically plausible models of electrophysiological responses. In other words, the macroscopic features of cortical activity can be modelled in terms of the microscopic behaviour of neurons. An evoked response potential (ERP) is the mean electrical potential measured from an electrode on the scalp, in response to some event. The purpose of this paper is to outline a population density approach to modelling ERPs. We propose a biologically plausible model of neuronal activity that enables the estimation of physiologically meaningful parameters from electrophysiological data. The model encompasses four basic characteristics of neuronal activity and organization: (i) neurons are dynamic units, (ii) driven by stochastic forces, (iii) organized into populations with similar biophysical properties and response characteristics and (iv) multiple populations interact to form functional networks. This leads to a formulation of population dynamics in terms of the Fokker-Planck equation. The solution of this equation is the temporal evolution of a probability density over state-space, representing the distribution of an ensemble of trajectories. Each trajectory corresponds to the changing state of a neuron. Measurements can be modelled by taking expectations over this density, e.g. mean membrane potential, firing rate or energy consumption per neuron. The key motivation behind our approach is that ERPs represent an average response over many neurons. This means it is sufficient to model the probability density over neurons, because this implicitly models their average state. Although the dynamics of each neuron can be highly stochastic, the dynamics of the density is not. This means we can use Bayesian inference and estimation tools that have already been established for deterministic systems. The potential importance of modelling density dynamics (as opposed to more conventional neural mass models) is that they include interactions among the moments of neuronal states (e.g. the mean depolarization may depend on the variance of synaptic currents through nonlinear mechanisms).Here, we formulate a population model, based on biologically informed model-neurons with spike-rate adaptation and synaptic dynamics. Neuronal sub-populations are coupled to form an observation model, with the aim of estimating and making inferences about coupling among sub-populations using real data. We approximate the time-dependent solution of the system using a bi-orthogonal set and first-order perturbation expansion. For didactic purposes, the model is developed first in the context of deterministic input, and then extended to include stochastic effects. The approach is demonstrated using synthetic data, where model parameters are identified using a Bayesian estimation scheme we have described previously.

Bayes Theorem↗

Demographic history of HIV-1 subtypes B and F in Brazil.

The reconstruction of the epidemic history of several HIV populations, by using methods that infer the population history from sampled gene sequence data, has revealed important subtype-specific and regional-specific differences in patterns of epidemic growth. Here, we employ Bayesian coalescent-based methods to compare the population history of the HIV-1 subtype B and F1 epidemics in Brazil from non-contemporary env and pol gene sequences. Our results suggest that after the introduction of the subtypes B and F1 into Brazilian population, around mid to late 1960s and late 1970s, respectively, these subtypes experienced an initial period of exponential growth with similar epidemic growth rates ( approximately 0.5-0.6year(-1)). Later, the spreading rate of both subtypes seems to have slowed-down since mid to late 1980s. This demographic pattern is very similar to that reported for the subtype B epidemics in high-income countries where HIV was initially transmitted through homosexual intercourse and injecting drug use, as in Brazil; suggesting that the characteristics of transmission networks may be a key determinant of the HIV epidemic growth pattern. It is important to note that most of the subtype B and F1 sequences used in this study come from the Southeast region that has been the most affected by the AIDS epidemic in Brazil, being responsible for around 63% of all AIDS cases reported since the early eighties; but may not represent the demographic trend of the HIV-1 epidemic in other Brazilian regions.

Bayes Theorem↗

Advances on BYY harmony learning: information theoretic perspective, generalized projection geometry, and independent factor autodetermination.

The nature of Bayesian Ying-Yang harmony learning is reexamined from an information theoretic perspective. Not only its ability for model selection and regularization is explained with new insights, but also discussions are made on its relations and differences from the studies of minimum description length (MDL), Bayesian approach, the bit-back based MDL, Akaike information criterion (AIC), maximum likelihood, information geometry, Helmholtz machines, and variational approximation. Moreover, a generalized projection geometry is introduced for further understanding such a new mechanism. Furthermore, new algorithms are also developed for implementing Gaussian factor analysis (FA) and non-Gaussian factor analysis (NFA) such that selecting appropriate factors is automatically made during parameter learning.

Algorithms↗

Gradient descent learning in and out of equilibrium.

Relations between the off thermal equilibrium dynamical process of on-line learning and the thermally equilibrated off-line learning are studied for potential gradient descent learning. The approach of Opper to study on-line Bayesian algorithms is used for potential based or maximum likelihood learning. We look at the on-line learning algorithm that best approximates the off-line algorithm in the sense of least Kullback-Leibler information loss. The closest on-line algorithm works by updating the weights along the gradient of an effective potential, which is different from the parent off-line potential. A few examples are analyzed and the origin of the potential annealing is discussed.

Algorithms↗

Selecting optimal experiments for multiple output multilayer perceptrons.

Where should a researcher conduct experiments to provide training data for a multilayer perceptron? This question is investigated, and a statistical method for selecting optimal experimental design points for multiple output multilayer perceptrons is introduced. Multiple class discrimination problems are examined using a framework in which the multilayer perceptron is viewed as a multivariate nonlinear regression model. Following a Bayesian formulation for the case where the variance-covariance matrix of the responses is unknown, a selection criterion is developed. This criterion is based on the volume of the joint confidence ellipsoid for the weights in a multilayer perceptron. An example is used to demonstrate the superiority of optimally selected design points over randomly chosen points, as well as points chosen in a grid pattern. Simplification of the basic criterion is offered through the use of Hadamard matrices to produce uncorrelated outputs.

Algorithms↗

Diversity of model approaches for breast cancer screening: a review of model assumptions by the Cancer Intervention and Surveillance Network (CISNET) Breast Cancer Groups.

The National Cancer Institute-sponsored Cancer Intervention and Surveillance Network program on breast cancer is composed of seven research groups working largely independently to model the impact of screening and adjuvant therapy on breast cancer mortality trends in the US from 1975 to 2000. Each of the groups has chosen a different modeling methodology without purposeful attempt to be in contrast with each other. The seven groups have met biannually since November 2000 to discuss their methodology and results. This article investigates the differences in methodology. To facilitate this comparison, each of the groups submitted a description of their model into a uniformly structured web based 'model profiler'. Six of the seven models simulate a preclinical natural history that cannot be observed directly with parameters estimated from published evidence concerning screening and therapy effects. The remaining model regards published evidence on intervention effects as prior information and updates that with information from the US population in a Bayesian type analysis. In general, the differences between the models appear to be small, particularly among the models driven by natural history assumptions. However, we demonstrate that such apparently small differences can have a large impact on surveillance of population trends. We describe a systematic approach to evaluating differences in model assumptions and results, as well as differences in modeling culture underlying the differences in model structure and parameters.

Breast Neoplasms↗

Bayesian model search for mixture models based on optimizing variational bounds.

When learning a mixture model, we suffer from the local optima and model structure determination problems. In this paper, we present a method for simultaneously solving these problems based on the variational Bayesian (VB) framework. First, in the VB framework, we derive an objective function that can simultaneously optimize both model parameter distributions and model structure. Next, focusing on mixture models, we present a deterministic algorithm to approximately optimize the objective function by using the idea of the split and merge operations which we previously proposed within the maximum likelihood framework. Then, we apply the method to mixture of expers (MoE) models to experimentally show that the proposed method can find the optimal number of experts of a MoE while avoiding local maxima.

Algorithms↗

Interpreting neuronal population activity by reconstruction: unified framework with application to hippocampal place cells.

Physical variables such as the orientation of a line in the visual field or the location of the body in space are coded as activity levels in populations of neurons. Reconstruction or decoding is an inverse problem in which the physical variables are estimated from observed neural activity. Reconstruction is useful first in quantifying how much information about the physical variables is present in the population and, second, in providing insight into how the brain might use distributed representations in solving related computational problems such as visual object recognition and spatial navigation. Two classes of reconstruction methods, namely, probabilistic or Bayesian methods and basis function methods, are discussed. They include important existing methods as special cases, such as population vector coding, optimal linear estimation, and template matching. As a representative example for the reconstruction problem, different methods were applied to multi-electrode spike train data from hippocampal place cells in freely moving rats. The reconstruction accuracy of the trajectories of the rats was compared for the different methods. Bayesian methods were especially accurate when a continuity constraint was enforced, and the best errors were within a factor of two of the information-theoretic limit on how accurate any reconstruction can be and were comparable with the intrinsic experimental errors in position tracking. In addition, the reconstruction analysis uncovered some interesting aspects of place cell activity, such as the tendency for erratic jumps of the reconstructed trajectory when the animal stopped running. In general, the theoretical values of the minimal achievable reconstruction errors quantify how accurately a physical variable is encoded in the neuronal population in the sense of mean square error, regardless of the method used for reading out the information. One related result is that the theoretical accuracy is independent of the width of the Gaussian tuning function only in two dimensions. Finally, all the reconstruction methods considered in this paper can be implemented by a unified neural network architecture, which the brain feasibly could use to solve related problems.

Action Potentials↗