Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian networks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Quantification of atherosclerotic plaque components using in vivo MRI and supervised classifiers.

In this work we aimed to study the possibility of using supervised classifiers to quantify the main components of carotid atherosclerotic plaque in vivo on the basis of multisequence MRI data. MRI data consisting of five MR weightings were obtained from 25 symptomatic subjects. Histological micrographs of endarterectomy specimens from the 25 carotids were used as a standard of reference for training and evaluation. The set of subjects was divided in a training set (12 subjects) and an evaluation set (13 subjects). Four different classifiers and two human MRI readers determined the percentages of calcified tissue, fibrous tissue, lipid core, and intraplaque hemorrhage on the subject level for all subjects in the evaluation set. Quantification of the relatively small amounts of calcium could not be done with statistical significance by either the classifiers or the MRI readers. For the other tissues a simple Bayesian classifier (Bayes) performed better than the other classifiers and the MRI readers. All classifiers performed better than the MRI readers in quantifying the sum of hemorrhage and lipid proportions. The MRI readers overestimated the hemorrhage proportions and tended to underestimate the lipid proportions. In conclusion, this pilot study demonstrates the benefits of algorithmic classifiers for quantifying plaque components.

Algorithms↗

A Bayesian framework for understanding texture segmentation in the primary visual cortex.

This paper presents a mathematical theory for understanding the computations involved in texture segmentation in the primary visual cortex. We propose that texture segmentation is a part of the early visual system's overall strategy to infer surfaces of objects in a visual scene. Based on this insight, we use the Bayesian inference paradigm to formulate the texture segmentation problem into a maximum a posteriori surface inference problem. The dynamical system for finding the optimal solution of this problem can be characterized by two concurrent and interactive processes: a gradual sharpening of the boundary signals and a simultaneous smoothing of the surface signals. The behavior of these dynamical processes was studied using both analytical and computational methods. We present some computational results and mathematical predictions. This theory suggests a novel framework for understanding the functional roles of the complex cells in the primary visual cortex.

Bayes Theorem↗

Clustering of genes into regulons using integrated modeling-COGRIM.

We present a Bayesian hierarchical model and Gibbs Sampling implementation that integrates gene expression, ChIP binding, and transcription factor motif data in a principled and robust fashion. COGRIM was applied to both unicellular and mammalian organisms under different scenarios of available data. In these applications, we demonstrate the ability to predict gene-transcription factor interactions with reduced numbers of false-positive findings and to make predictions beyond what is obtained when single types of data are considered.

CCAAT-Enhancer-Binding Protein-beta↗

Spatio-temporal autoregressive models defined over brain manifolds.

Multivariate Autoregressive time series models (MAR) are an increasingly used tool for exploring functional connectivity in Neuroimaging. They provide the framework for analyzing the Granger Causality of a given brain region on others. In this article, we shall limit our attention to linear MAR models, in which a set of matrices of autoregressive coefficients Ak (k = 1,...,p) describe the dependence of present values of the image on lagged values of its past. Methods for estimating the Ak and determining which elements that are zero are well-known and are the basis for directed measures of influence. However, to date, MAR models are limited in the number of time series they can handle, forcing the a priori selection of a (small) number of voxels or regions of interest for analysis. This ignores the full spatio-temporal nature of functional brain data which are, in fact, collections of time series sampled over an underlying continuous spatial manifold the brain. A fully spatio-temporal MAR model (ST-MAR) is developed within the framework of functional data analysis. For spatial data, each row of a matrix Ak is the influence field of a given voxel. A Bayesian ST-MAR model is specified in which the influence fields for all voxels are required to vary smoothly over space. This requirement is enforced by penalizing the spatial roughness of the influence fields. This roughness is calculated with a discrete version of the spatial Laplacian operator. A massive reduction in dimensionality of computations is achieved via the singular value decomposition, making an interactive exploration of the model feasible. Use of the model is illustrated with an fMRI time series that was gathered concurrently with EEG in order to analyze the origin of resting brain rhythms.

Bayes Theorem↗

Efficient information theoretic strategies for classifier combination, feature extraction and performance evaluation in improving false positives and false negatives for spam e-mail filtering.

Spam emails are considered as a serious privacy-related violation, besides being a costly, unsolicited communication. Various spam filtering techniques have been so far proposed, mainly based on Naïve Bayesian algorithms. Other Machine Learning algorithms like Boosting trees, or Support Vector Machines (SVM) have already been used with success. However, the number of False Positives (FP) and False Negatives (FN) resulting through applying various spam e-mail filters still remains too high and the problem of spam e-mail categorization cannot be solved completely from a practical viewpoint. In this paper, we propose a novel approach for spam e-mail filtering based on efficient information theoretic techniques for integrating classifiers, for extracting improved features and for properly evaluating categorization accuracy in terms of FP and FN. The goal of the presented methodology is to empirically but explicitly minimize these FP and FN numbers by combining high-performance FP filters with high-performance FN filters emerging from a previous work of the authors [Zorkadis, V., Panayotou, M., & Karras, D. A. (2005). Improved spam e-mail filtering based on committee machines and information theoretic feature extraction. Proceedings of the International Joint Conference on Neural Networks, July 31-August 4, 2005, Montreal, Canada]. To this end, Random Committee-based filters along with ADTree-based ones are efficiently combined through information theory, respectively. The experiments conducted are of the most extensive ones so far in the literature, exploiting widely accepted benchmarking e-mail data sets and comparing the proposed methodology with the Naive Bayes spam filter as well as with the Boosting tree methodology, the classification via regression and other machine learning models. It is illustrated by means of novel information theoretic measures of FP & FN filtering performance that the proposed approach is very favorably compared to the other rival methods. Finally, it is found that the proposed information theoretic Boolean features present a remarkably high spam categorization performance.

Algorithms↗

Bayesian population decoding of motor cortical activity using a Kalman filter.

Effective neural motor prostheses require a method for decoding neural activity representing desired movement. In particular, the accurate reconstruction of a continuous motion signal is necessary for the control of devices such as computer cursors, robots, or a patient's own paralyzed limbs. For such applications, we developed a real-time system that uses Bayesian inference techniques to estimate hand motion from the firing rates of multiple neurons. In this study, we used recordings that were previously made in the arm area of primary motor cortex in awake behaving monkeys using a chronically implanted multielectrode microarray. Bayesian inference involves computing the posterior probability of the hand motion conditioned on a sequence of observed firing rates; this is formulated in terms of the product of a likelihood and a prior. The likelihood term models the probability of firing rates given a particular hand motion. We found that a linear gaussian model could be used to approximate this likelihood and could be readily learned from a small amount of training data. The prior term defines a probabilistic model of hand kinematics and was also taken to be a linear gaussian model. Decoding was performed using a Kalman filter, which gives an efficient recursive method for Bayesian inference when the likelihood and prior are linear and gaussian. In off-line experiments, the Kalman filter reconstructions of hand trajectory were more accurate than previously reported results. The resulting decoding algorithm provides a principled probabilistic model of motor-cortical coding, decodes hand motion in real time, provides an estimate of uncertainty, and is straightforward to implement. Additionally the formulation unifies and extends previous models of neural coding while providing insights into the motor-cortical code.

Action Potentials↗

Statistical methods for linking health, exposure, and hazards.

The Environmental Public Health Tracking Network (EPHTN) proposes to link environmental hazards and exposures to health outcomes. Statistical methods used in case-control and cohort studies to link health outcomes to individual exposure estimates are well developed. However, reliable exposure estimates for many contaminants are not available at the individual level. In these cases, exposure/hazard data are often aggregated over a geographic area, and ecologic models are used to relate health outcome and exposure/hazard. Ecologic models are not without limitations in interpretation. EPHTN data are characteristic of much information currently being collected--they are multivariate, with many predictors and response variables, often aggregated over geographic regions (small and large) and correlated in space and/or time. The methods to model trends in space and time, handle correlation structures in the data, estimate effects, test hypotheses, and predict future outcomes are relatively new and without extensive application in environmental public health. In this article we outline a tiered approach to data analysis for EPHTN and review the use of standard methods for relating exposure/hazards, disease mapping and clustering techniques, Bayesian approaches, Markov chain Monte Carlo methods for estimation of posterior parameters, and geostatistical methods. The advantages and limitations of these methods are discussed.

Case-Control Studies↗

A Bayesian method for identifying missing enzymes in predicted metabolic pathway databases.

BACKGROUND: The PathoLogic program constructs Pathway/Genome databases by using a genome's annotation to predict the set of metabolic pathways present in an organism. PathoLogic determines the set of reactions composing those pathways from the enzymes annotated in the organism's genome. Most annotation efforts fail to assign function to 40-60% of sequences. In addition, large numbers of sequences may have non-specific annotations (e.g., thiolase family protein). Pathway holes occur when a genome appears to lack the enzymes needed to catalyze reactions in a pathway. If a protein has not been assigned a specific function during the annotation process, any reaction catalyzed by that protein will appear as a missing enzyme or pathway hole in a Pathway/Genome database. RESULTS: We have developed a method that efficiently combines homology and pathway-based evidence to identify candidates for filling pathway holes in Pathway/Genome databases. Our program not only identifies potential candidate sequences for pathway holes, but combines data from multiple, heterogeneous sources to assess the likelihood that a candidate has the required function. Our algorithm emulates the manual sequence annotation process, considering not only evidence from homology searches, but also considering evidence from genomic context (i.e., is the gene part of an operon?) and functional context (e.g., are there functionally-related genes nearby in the genome?) to determine the posterior belief that a candidate has the required function. The method can be applied across an entire metabolic pathway network and is generally applicable to any pathway database. The program uses a set of sequences encoding the required activity in other genomes to identify candidate proteins in the genome of interest, and then evaluates each candidate by using a simple Bayes classifier to determine the probability that the candidate has the desired function. We achieved 71% precision at a probability threshold of 0.9 during cross-validation using known reactions in computationally-predicted pathway databases. After applying our method to 513 pathway holes in 333 pathways from three Pathway/Genome databases, we increased the number of complete pathways by 42%. We made putative assignments to 46% of the holes, including annotation of 17 sequences of previously unknown function. CONCLUSIONS: Our pathway hole filler can be used not only to increase the utility of Pathway/Genome databases to both experimental and computational researchers, but also to improve predictions of protein function.

Amino Acid Oxidoreductases↗

The context-tree kernel for strings.

We propose a new kernel for strings which borrows ideas and techniques from information theory and data compression. This kernel can be used in combination with any kernel method, in particular Support Vector Machines for string classification, with notable applications in proteomics. By using a Bayesian averaging framework with conjugate priors on a class of Markovian models known as probabilistic suffix trees or context-trees, we compute the value of this kernel in linear time and space while only using the information contained in the spectrum of the considered strings. This is ensured through an adaptation of a compression method known as the context-tree weighting algorithm. Encouraging classification results are reported on a standard protein homology detection experiment, showing that the context-tree kernel performs well with respect to other state-of-the-art methods while using no biological prior knowledge.

Algorithms↗

Transmission history of major China-prevalent Mycobacterium tuberculosis sub-lineages in East Asia.

Mycobacterium tuberculosis complex (MTBC) is distributed globally and has posed a severe threat to human health throughout history. In this study, we analyzed whole-genome data from the four major MTBC sub-lineages prevalent in China (L2.2, L4.2, L4.4, and L4.5) to reconstruct their transmission and expansion histories across East Asia and parts of Central Asia. We found that L2.2 has established a highly connected transmission network centered in Southern China, whereas L4.2 is characterized by cross-border transmission between Central Asia and Western China, and L4.4 and L4.5 exhibit repeated transmission events between Southeast Asia and Southern China. By reconstructing their population histories, we demonstrated that these sub-lineages have experienced multi-stage expansions since the 15th century, accompanied by a recent rapid proliferation of evolutionary clades. These findings reveal that the MTBC epidemic in East Asia may follow a pattern of long-term historical adaptation superimposed with recent concentrated outbreaks, providing potential genomic evidence to inform precise regional tuberculosis control strategies in China.

Mycobacterium tuberculosis↗

Bayesian approach to discovering pathogenic SNPs in conserved protein domains.

The success rate of association studies can be improved by selecting better genetic markers for genotyping or by providing better leads for identifying pathogenic single nucleotide polymorphisms (SNPs) in the regions of linkage disequilibrium with positive disease associations. We have developed a novel algorithm to predict pathogenic single amino acid changes, either nonsynonymous SNPs (nsSNPs) or missense mutations, in conserved protein domains. Using a Bayesian framework, we found that the probability of a microbial missense mutation causing a significant change in phenotype depended on how much difference it made in several phylogenetic, biochemical, and structural features related to the single amino acid substitution. We tested our model on pathogenic allelic variants (missense mutations or nsSNPs) included in OMIM, and on the other nsSNPs in the same genes (from dbSNP) as the nonpathogenic variants. As a result, our model predicted pathogenic variants with a 10% false-positive rate. The high specificity of our prediction algorithm should make it valuable in genetic association studies aimed at identifying pathogenic SNPs.

Algorithms↗

The RIN: an RNA integrity number for assigning integrity values to RNA measurements.

BACKGROUND: The integrity of RNA molecules is of paramount importance for experiments that try to reflect the snapshot of gene expression at the moment of RNA extraction. Until recently, there has been no reliable standard for estimating the integrity of RNA samples and the ratio of 28S:18S ribosomal RNA, the common measure for this purpose, has been shown to be inconsistent. The advent of microcapillary electrophoretic RNA separation provides the basis for an automated high-throughput approach, in order to estimate the integrity of RNA samples in an unambiguous way. METHODS: A method is introduced that automatically selects features from signal measurements and constructs regression models based on a Bayesian learning technique. Feature spaces of different dimensionality are compared in the Bayesian framework, which allows selecting a final feature combination corresponding to models with high posterior probability. RESULTS: This approach is applied to a large collection of electrophoretic RNA measurements recorded with an Agilent 2100 bioanalyzer to extract an algorithm that describes RNA integrity. The resulting algorithm is a user-independent, automated and reliable procedure for standardization of RNA quality control that allows the calculation of an RNA integrity number (RIN). CONCLUSION: Our results show the importance of taking characteristics of several regions of the recorded electropherogram into account in order to get a robust and reliable prediction of RNA integrity, especially if compared to traditional methods.

Algorithms↗

A Bayesian molecular interaction library.

We describe a library of molecular fragments designed to model and predict non-bonded interactions between atoms. We apply the Bayesian approach, whereby prior knowledge and uncertainty of the mathematical model are incorporated into the estimated model and its parameters. The molecular interaction data are strengthened by narrowing the atom classification to 14 atom types, focusing on independent molecular contacts that lie within a short cutoff distance, and symmetrizing the interaction data for the molecular fragments. Furthermore, the location of atoms in contact with a molecular fragment are modeled by Gaussian mixture densities whose maximum a posteriori estimates are obtained by applying a version of the expectation-maximization algorithm that incorporates hyperparameters for the components of the Gaussian mixtures. A routine is introduced providing the hyperparameters and the initial values of the parameters of the Gaussian mixture densities. A model selection criterion, based on the concept of a 'minimum message length' is used to automatically select the optimal complexity of a mixture model and the most suitable orientation of a reference frame for a fragment in a coordinate system. The type of atom interacting with a molecular fragment is predicted by values of the posterior probability function and the accuracy of these predictions is evaluated by comparing the predicted atom type with the actual atom type seen in crystal structures. The fact that an atom will simultaneously interact with several molecular fragments forming a cohesive network of interactions is exploited by introducing two strategies that combine the predictions of atom types given by multiple fragments. The accuracy of these combined predictions is compared with those based on an individual fragment. Exhaustive validation analyses and qualitative examples (e.g., the ligand-binding domain of glutamate receptors) demonstrate that these improvements lead to effective modeling and prediction of molecular interactions.

Algorithms↗

[An unusual excess of mortality in a small Tuscan municipality and the "nursing home effect"].

OBJECTIVE: In the last decades unusual mortality excesses were observed in the small area of Montaione, where the main activities are agriculture and tourism. The aim of this study was to evaluate if the observed excess mortality had to be attributed to the deaths occurred among the local large Nursing Home's guests which were half of the total deaths registered among residents. DESIGN: Empirical Bayesian Mortality Ratios (EBMR), applying the method of Clayton and Kaldor, were calculated either including or excluding the guests of the Nursing Home from the deaths and from the population. Only the population with age > or = 65 was included in the analysis. The expected deaths were calculated using the Tuscan population mortality rates by sex, age and specific cause of death (all causes, cardiovascular diseases, cerebrovascular diseases, digestive diseases and respiratory diseases). RESULTS: Excluding the guests of the Nursing Home from the analysis it was observed a strong decrease of the EBMRs for almost all causes considered, but those for cerebrovascular and respiratory diseases. CONCLUSION: The results obtained underline the necessity to take in consideration also a possible "Nursing Home effect" in evaluating mortality excesses in small areas.

Aged↗

Implementing the Bayesian paradigm: reporting research results over the World-Wide Web.

For decades, statisticians, philosophers, medical investigators and others interested in data analysis have argued that the Bayesian paradigm is the proper approach for reporting the results of scientific analyses for use by clients and readers. To date, the methods have been too complicated for non-statisticians to use. In this paper we argue that the World-Wide Web provides the perfect environment to put the Bayesian paradigm into practice: the likelihood function of the data is parsimoniously represented on the server side, the reader uses the client to represent her prior belief, and a downloaded program (a Java applet) performs the combination. In our approach, a different applet can be used for each likelihood function, prior belief can be assessed graphically, and calculation results can be reported in a variety of ways. We present a prototype implementation, BayesApplet, for two-arm clinical trials with normally-distributed outcomes, a prominent model for clinical trials. The primary implication of this work is that publishing medical research results on the Web can take a form beyond or different from that currently used on paper, and can have a profound impact on the publication and use of research results.

Bayes Theorem↗

Accurate and fast off and online fuzzy ARTMAP-based image classification with application to genetic abnormality diagnosis.

We propose and investigate the fuzzy ARTMAP neural network in off and online classification of fluorescence in situ hybridization image signals enabling clinical diagnosis of numerical genetic abnormalities. We evaluate the classification task (detecting a several abnormalities separately or simultaneously), classifier paradigm (monolithic or hierarchical), ordering strategy for the training patterns (averaging or voting), training mode (for one epoch, with validation or until completion) and model sensitivity to parameters. We find the fuzzy ARTMAP accurate in accomplishing both tasks requiring only very few training epochs. Also, selecting a training ordering by voting is more precise than if averaging over orderings. If trained for only one epoch, the fuzzy ARTMAP provides fast, yet stable and accurate learning as well as insensitivity to model complexity. Early stop of training using a validation set reduces the fuzzy ARTMAP complexity as for other machine learning models but cannot improve accuracy beyond that achieved when training is completed. Compared to other machine learning models, the fuzzy ARTMAP does not loose but gain accuracy when overtrained, although increasing its number of categories. Learned incrementally, the fuzzy ARTMAP reaches its ultimate accuracy very fast obtaining most of its data representation capability and accuracy by using only a few examples. Finally, the fuzzy ARTMAP accuracy for this domain is comparable with those of the multilayer perceptron and support vector machine and superior to those of the naive Bayesian and linear classifiers.

Artificial Intelligence↗

Parallel Metropolis coupled Markov chain Monte Carlo for Bayesian phylogenetic inference.

MOTIVATION: Bayesian estimation of phylogeny is based on the posterior probability distribution of trees. Currently, the only numerical method that can effectively approximate posterior probabilities of trees is Markov chain Monte Carlo (MCMC). Standard implementations of MCMC can be prone to entrapment in local optima. Metropolis coupled MCMC [(MC)(3)], a variant of MCMC, allows multiple peaks in the landscape of trees to be more readily explored, but at the cost of increased execution time. RESULTS: This paper presents a parallel algorithm for (MC)(3). The proposed parallel algorithm retains the ability to explore multiple peaks in the posterior distribution of trees while maintaining a fast execution time. The algorithm has been implemented using two popular parallel programming models: message passing and shared memory. Performance results indicate nearly linear speed improvement in both programming models for small and large data sets.

Algorithms↗

Methods for combining experts' probability assessments.

This article reviews statistical techniques for combining multiple probability distributions. The framework is that of a decision maker who consults several experts regarding some events. The experts express their opinions in the form of probability distributions. The decision maker must aggregate the experts' distributions into a single distribution that can be used for decision making. Two classes of aggregation methods are reviewed. When using a supra Bayesian procedure, the decision maker treats the expert opinions as data that may be combined with its own prior distribution via Bayes' rule. When using a linear opinion pool, the decision maker forms a linear combination of the expert opinions. The major feature that makes the aggregation of expert opinions difficult is the high correlation or dependence that typically occurs among these opinions. A theme of this paper is the need for training procedures that result in experts with relatively independent opinions or for aggregation methods that implicitly or explicitly model the dependence among the experts. Analyses are presented that show that m dependent experts are worth the same as k independent experts where k < or = m. In some cases, an exact value for k can be given; in other cases, lower and upper bounds can be placed on k.

Bayes Theorem↗