Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian networks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Predicting gene regulation by sigma factors in Bacillus subtilis from genome-wide data.

MOTIVATION: Sigma factors regulate the expression of genes in Bacillus subtilis at the transcriptional level. We assess the accuracy of a fold-change analysis, Bayesian networks, dynamic models and supervised learning based on coregulation in predicting gene regulation by sigma factors from gene expression data. To improve the prediction accuracy, we combine sequence information with expression data by adding their log-likelihood scores and by using a logistic regression model. We use the resulting score function to discover currently unknown gene regulations by sigma factors. RESULTS: The coregulation-based supervised learning method gave the most accurate prediction of sigma factors from expression data. We found that the logistic regression model effectively combines expression data with sequence information. In a genome-wide search, highly significant logistic regression scores were found for several genes whose transcriptional regulation is currently unknown. We provide the corresponding RNA polymerase binding sites to enable a straightforward experimental verification of these predictions.

Algorithms↗

Analyzing microarray data using quantitative association rules.

MOTIVATION: We tackle the problem of finding regularities in microarray data. Various data mining tools, such as clustering, classification, Bayesian networks and association rules, have been applied so far to gain insight into gene-expression data. Association rule mining techniques used so far work on discretizations of the data and cannot account for cumulative effects. In this paper, we investigate the use of quantitative association rules that can operate directly on numeric data and represent cumulative effects of variables. Technically speaking, this type of quantitative association rules based on half-spaces can find non-axis-parallel regularities. RESULTS: We performed a variety of experiments testing the utility of quantitative association rules for microarray data. First of all, the results should be statistically significant and robust against fluctuations in the data. Next, the approach should be scalable in the number of variables, which is important for such high-dimensional data. Finally, the rules should make sense biologically and be sufficiently different from rules found in regular association rule mining working with discretizations. In all of these dimensions, the proposed approach performed satisfactorily. Therefore, quantitative association rules based on half-spaces should be considered as a tool for the analysis of microarray gene-expression data. AVAILABILITY: The code is available from the authors on request.

Algorithms↗

BIAS: Bioinformatics Integrated Application Software.

MOTIVATION: We introduce a development platform especially tailored to Bioinformatics research and software development. BIAS (Bioinformatics Integrated Application Software) provides the tools necessary for carrying out integrative Bioinformatics research requiring multiple datasets and analysis tools. It follows an object-relational strategy for providing persistent objects, allows third-party tools to be easily incorporated within the system and supports standards and data-exchange protocols common to Bioinformatics. AVAILABILITY: BIAS is an OpenSource project and is freely available to all interested users at http://www.mcb.mcgill.ca/~bias/. This website also contains a paper containing a more detailed description of BIAS and a sample implementation of a Bayesian network approach for the simultaneous prediction of gene regulation events and of mRNA expression from combinations of gene regulation events. CONTACT: hallett@mcb.mcgill.ca.

Computational Biology↗

A multiple-feature framework for modelling and predicting transcription factor binding sites.

MOTIVATION: The identification of transcription factor binding sites in promoter sequences is an important problem, since it reveals information about the transcriptional regulation of genes. For analysing transcriptional regulation, computational approaches for predicting putative binding sites are applied. Commonly used stochastic models for binding sites are position-specific score matrices, which show weak predictive power. RESULTS: We have developed a probabilistic modelling approach, which allows to consider diverse characteristic binding site properties to obtain more accurate representations of binding sites. These properties are modelled as random variables in Bayesian networks, which are capable of dealing with dependencies among binding site properties. Cross-validation on several datasets shows improvements in the false positive error rate and the significance (P-value) of true binding sites.

Algorithms↗

Dragon Promoter Mapper (DPM): a Bayesian framework for modelling promoter structures.

UNLABELLED: Dragon Promoter Mapper (DPM) is a tool to model promoter structure of co-regulated genes using methodology of Bayesian networks. DPM exploits an exhaustive set of motif features (such as motif, its strand, the order of motif occurrence and mutual distance between the adjacent motifs) and generates models from the target promoter sequences, which may be used to (1) detect regions in a genomic sequence which are similar to the target promoters or (2) to classify other promoters as similar or not to the target promoter group. DPM can also be used for modelling of enhancers and silencers. AVAILABILITY: http://defiant.i2r.a-star.edu.sg/projects/BayesPromoter/ CONTACT: vlad@sanbi.ac.za SUPPLEMENTARY INFORMATION: Manual for using DPM web server is provided at http://defiant.i2r.a-star.edu.sg/projects/BayesPromoter/html/manual/manual.htm.

Algorithms↗

A framework of integrating gene relations from heterogeneous data sources: an experiment on Arabidopsis thaliana.

One of the most important goals of biological investigation is to uncover gene functional relations. In this study we propose a framework for extraction and integration of gene functional relations from diverse biological data sources, including gene expression data, biological literature and genomic sequence information. We introduce a two-layered Bayesian network approach to integrate relations from multiple sources into a genome-wide functional network. An experimental study was conducted on a test-bed of Arabidopsis thaliana. Evaluation of the integrated network demonstrated that relation integration could improve the reliability of relations by combining evidence from different data sources. Domain expert judgments on the gene functional clusters in the network confirmed the validity of our approach for relation integration and network inference.

Arabidopsis↗

Protease mutation M89I/V is linked to therapy failure in patients infected with the HIV-1 non-B subtypes C, F or G.

OBJECTIVE: To investigate whether and how mutations at position 89 of HIV-1 protease were associated with protease inhibitor (PI) failure, and what is the impact of the HIV-1 subtype. METHODS: In a database containing pol nucleotide sequences and treatment history, the correlation between PI experience and mutations at codon 89 was determined separately for subtype B and several non-B subtypes. A Bayesian network model was used to map the resistance pathways in which M89I/V is involved for subtype G. The phenotypic effect of M89I/V for several PIs was also measured. RESULTS: The analysis showed that for the subtypes C, F and G in which the wild-type codon at 89 was M compared to L for subtype B, M89I/V was significantly more frequently observed in PI-treated patients displaying major resistance mutations to PIs than in drug-naive patients. M89I/V was strongly associated with PI resistance mutations at codons 71, 74 and 90. Phenotypically, M89I/V alone did not confer a reduced susceptibility to PIs. However, when combined with L90M, a significantly reduced susceptibility to nelfinavir was observed (P < 0.05) in comparison with strains with L90M alone. CONCLUSIONS: The results of the present study show that M89I/V is associated with PI experience in subtypes C, F and G but not in subtype B. M89I/V should be considered a secondary PI mutation with an important effect on nelfinavir susceptibility in the presence of L90M.

Adult↗

Bayesian analysis, pattern analysis, and data mining in health care.

PURPOSE OF REVIEW: To discuss the current role of data mining and Bayesian methods in biomedicine and heath care, in particular critical care. RECENT FINDINGS: Bayesian networks and other probabilistic graphical models are beginning to emerge as methods for discovering patterns in biomedical data and also as a basis for the representation of the uncertainties underlying clinical decision-making. At the same time, techniques from machine learning are being used to solve biomedical and health-care problems. SUMMARY: With the increasing availability of biomedical and health-care data with a wide range of characteristics there is an increasing need to use methods which allow modeling the uncertainties that come with the problem, are capable of dealing with missing data, allow integrating data from various sources, explicitly indicate statistical dependence and independence, and allow integrating biomedical and clinical background knowledge. These requirements have given rise to an influx of new methods into the field of data analysis in health care, in particular from the fields of machine learning and probabilistic graphical models.

Bayes Theorem↗

Predicting subcellular localization via protein motif co-occurrence.

The prediction of subcellular localization of proteins from their primary sequence is a challenging problem in bioinformatics. We have created a Bayesian network localization predictor called PSLT that is based on the combinatorial presence of InterPro motifs and specific membrane domains in human proteins. This probabilistic framework generates a likelihood of localization to all organelles and allows to predict multicompartmental proteins. When used to predict on nine compartments, PSLT achieves an accuracy of 78% as estimated by using a 10-fold cross-validation test and a coverage of 74%. When used to predict the localization of proteins from other closely related species, it achieves a prediction accuracy and a coverage >80%. We compared the localization predictions of PSLT to those determined through GFP-tagging and microscopy for a group of human proteins. We found two general classes of proteins that are mislocalized by the GFP-tagging strategy but are correctly localized by PSLT. This suggests that PSLT can be used in combination with experimental approaches for localization to identify proteins for which additional experimental validation is required. We used our predictor to annotate all 9793 human proteins from SWISS-PROT release 41.25, 16% of which are predicted by PSLT to be present in more than one compartment.

Humans↗

Predicting gene function from patterns of annotation.

The Gene Ontology (GO) Consortium has produced a controlled vocabulary for annotation of gene function that is used in many organism-specific gene annotation databases. This allows the prediction of gene function based on patterns of annotation. For example, if annotations for two attributes tend to occur together in a database, then a gene holding one attribute is likely to hold the other as well. We modeled the relationships among GO attributes with decision trees and Bayesian networks, using the annotations in the Saccharomyces Genome Database (SGD) and in FlyBase as training data. We tested the models using cross-validation, and we manually assessed 100 gene-attribute associations that were predicted by the models but that were not present in the SGD or FlyBase databases. Of the 100 manually assessed associations, 41 were judged to be true, and another 42 were judged to be plausible.

Animals↗

Human-mouse gene identification by comparative evidence integration and evolutionary analysis.

The identification of genes in the human genome remains a challenge, as the actual predictions appear to disagree tremendously and vary dramatically on the basis of the specific gene-finding methodology used. Because the pattern of conservation in coding regions is expected to be different from intronic or intergenic regions, a comparative computational analysis can lead, in principle, to an improved computational identification of genes in the human genome by using a reference, such as mouse genome. However, this comparative methodology critically depends on three important factors: (1) the selection of the most appropriate reference genome. In particular, it is not clear whether the mouse is at the correct evolutionary distance from the human to provide sufficiently distinctive conservation levels in different genomic regions, (2) the selection of comparative features that provide the most benefit to gene recognition, and (3) the selection of evidence integration architecture that effectively interprets the comparative features. We address the first question by a novel evolutionary analysis that allows us to explicitly correlate the performance of the gene recognition system with the evolutionary distance (time) between the two genomes. Our simulation results indicate that there is a wide range of reference genomes at different evolutionary time points that appear to deliver reasonable comparative prediction of human genes. In particular, the evolutionary time between human and mouse generally falls in the region of good performance; however, better accuracy might be achieved with a reference genome further than mouse. To address the second question, we propose several natural comparative measures of conservation for identifying exons and exon boundaries. Finally, we experiment with Bayesian networks for the integration of comparative and compositional evidence.

Animals↗

Spatiotemporal video segmentation based on graphical models.

This paper proposes a probabilistic framework for spatiotemporal segmentation of video sequences. Motion information, boundary information from intensity segmentation, and spatial connectivity of segmentation are unified in the video segmentation process by means of graphical models. A Bayesian network is presented to model interactions among the motion vector field, the intensity segmentation field, and the video segmentation field. The notion of the Markov random field is used to encourage the formation of continuous regions. Given consecutive frames, the conditional joint probability density of the three fields is maximized in an iterative way. To effectively utilize boundary information from the intensity segmentation, distance transformation is employed in local objective functions. Experimental results show that the method is robust and generates spatiotemporally coherent segmentation results. Moreover, the proposed video segmentation approach can be viewed as the compromise of previous motion based approaches and region merging approaches.

Algorithms↗

"Shape activity": a continuous-state HMM for moving/deforming shapes with application to abnormal activity detection.

The aim is to model "activity" performed by a group of moving and interacting objects (which can be people, cars, or different rigid components of the human body) and use the models for abnormal activity detection. Previous approaches to modeling group activity include co-occurrence statistics (individual and joint histograms) and dynamic Bayesian networks, neither of which is applicable when the number of interacting objects is large. We treat the objects as point objects (referred to as "landmarks") and propose to model their changing configuration as a moving and deforming "shape" (using Kendall's shape theory for discrete landmarks). A continuous-state hidden Markov model is defined for landmark shape dynamics in an activity. The configuration of landmarks at a given time forms the observation vector, and the corresponding shape and the scaled Euclidean motion parameters form the hidden-state vector. An abnormal activity is then defined as a change in the shape activity model, which could be slow or drastic and whose parameters are unknown. Results are shown on a real abnormal activity-detection problem involving multiple moving objects.

Algorithms↗

Identifying the combination of genetic factors that determine susceptibility to cervical cancer.

Cervical cancer is common among women all over the world. Although infection with high-risk types of human papillomavirus (HPV) has been identified as the primary cause of cervical cancer, only some of those infected go on to develop cervical cancer. Obviously, the progression from HPV infection to cancer involves other environmental and host factors. Recent population-based twin and family studies have demonstrated the importance of the hereditary component of cervical cancer, associated with genetic susceptibility. Consequently, single-nucleotide polymorphism (SNP) markers and microsatellites should be considered genetic factors for determining what combinations of genetic factors are involved in precancerous changes to cervical cancer. This study employs a Bayesian network and four different decision tree algorithms, and compares the performance of these learning algorithms. The results of this study raise the possibility of investigations that could identify combinations of genetic factors, such as SNPs and microsatellites, that influence the risk associated with common complex multifactorial diseases, such as cervical cancer.

Algorithms↗

Modeling the health care costs of geriatric inpatients.

This paper extends a method for modeling the survival of patients in hospitals to allow the expected cost to be estimated for the patients' accumulated duration of time in care. An extension of Bayesian network (BN) theory has previously been developed to model patients' survival time in hospitals with respect to the graphical and probabilistic representation of the interrelationships between the patients' clinical variables. Unlike previous BN techniques, this extended model can accommodate continuous times that are skewed in nature. This paper presents the theory behind such an approach and extends it by attaching a cost variable to the survival times, enabling the costing and efficient management of groups of patients in hospitals. An application of the model is illustrated by considering a group of 4260 patients admitted into the geriatric department of a U.K. hospital between 1994-1997. Results are derived for the distribution for their length of stay in the hospital and associated costs. The model's practical use is highlighted by illustrating how hospital managers could benefit using such a method for investigating the influence of future decisions and policy changes on the hospital's expenditure.

Computer Simulation↗

Inferring figure-ground using a recurrent integrate-and-fire neural circuit.

Several theories of early visual perception hypothesize neural circuits that are responsible for assigning ownership of an object's occluding contour to a region which represents the "figure." Previously, we have presented a Bayesian network model which integrates multiple cues and uses belief propagation to infer local figure-ground relationships along an object's occluding contour. In this paper, we use a linear integrate-and-fire model to demonstrate how such inference mechanisms could be carried out in a biologically realistic neural circuit. The circuit maps the membrane potentials of individual neurons to log probabilities and uses recurrent connections to represent transition probabilities. The network's "perception" of figure-ground is demonstrated for several examples, including perceptually ambiguous figures, and compared qualitatively and quantitatively with human psychophysics.

Action Potentials↗

Semisupervised learning of classifiers: theory, algorithms, and their application to human-computer interaction.

Automatic classification is one of the basic tasks required in any pattern recognition and human computer interaction application. In this paper, we discuss training probabilistic classifiers with labeled and unlabeled data. We provide a new analysis that shows under what conditions unlabeled data can be used in learning to improve classification performance. We also show that, if the conditions are violated, using unlabeled data can be detrimental to classification performance. We discuss the implications of this analysis to a specific type of probabilistic classifiers, Bayesian networks, and propose a new structure learning algorithm that can utilize unlabeled data to improve classification. Finally, we show how the resulting algorithms are successfully employed in two applications related to human-computer interaction and pattern recognition: facial expression recognition and face detection.

Journal Article↗

Active and dynamic information fusion for facial expression understanding from image sequences.

This paper explores the use of multisensory information fusion technique with Dynamic Bayesian networks (DBNs) for modeling and understanding the temporal behaviors of facial expressions in image sequences. Our facial feature detection and tracking based on active IR illumination provides reliable visual information under variable lighting and head motion. Our approach to facial expression recognition lies in the proposed dynamic and probabilistic framework based on combining DBNs with Ekman's Facial Action Coding System (FACS) for systematically modeling the dynamic and stochastic behaviors of spontaneous facial expressions. The framework not only provides a coherent and unified hierarchical probabilistic framework to represent spatial and temporal information related to facial expressions, but also allows us to actively select the most informative visual cues from the available information sources to minimize the ambiguity in recognition. The recognition of facial expressions is accomplished by fusing not only from the current visual observations, but also from the previous visual evidences. Consequently, the recognition becomes more robust and accurate through explicitly modeling temporal behavior of facial expression. In this paper, we present the theoretical foundation underlying the proposed probabilistic and dynamic framework for facial expression modeling and understanding. Experimental results demonstrate that our approach can accurately and robustly recognize spontaneous facial expressions from an image sequence under different conditions.

Algorithms↗