Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73Linked to original sources

Inferring gene regulatory networks from time series data using the minimum description length principle.

MOTIVATION: A central question in reverse engineering of genetic networks consists in determining the dependencies and regulating relationships among genes. This paper addresses the problem of inferring genetic regulatory networks from time-series gene-expression profiles. By adopting a probabilistic modeling framework compatible with the family of models represented by dynamic Bayesian networks and probabilistic Boolean networks, this paper proposes a network inference algorithm to recover not only the direct gene connectivity but also the regulating orientations. RESULTS: Based on the minimum description length principle, a novel network inference algorithm is proposed that greatly shrinks the search space for graphical solutions and achieves a good trade-off between modeling complexity and data fitting. Simulation results show that the algorithm achieves good performance in the case of synthetic networks. Compared with existing state-of-the-art results in the literature, the proposed algorithm exceptionally excels in efficiency, accuracy, robustness and scalability. Given a time-series dataset for Drosophila melanogaster, the paper proposes a genetic regulatory network involved in Drosophila's muscle development. AVAILABILITY: Available from the authors upon request.

Algorithms↗

Effect of clinical mastitis on the lactation curve: a mixed model estimation using daily milk weights.

The objective of this study was to estimate the milk production losses associated with clinical mastitis using mixed linear models and correlation structures that have not been available previously. Data used included computer-recorded daily milk yields and detailed and accurate recordings of clinical mastitis cases. Two commercial Holstein dairy farms in New York State participated in the study, one with 650 lactating cows and another that began the study with 830 lactating cows and increased to 1120 cows by the end of the study. Cows on both farms were housed in free stall barns and milked 3 times daily in milking parlors. Electrical conductivity was used as a diagnostic aid for clinical mastitis on both farms. Date of clinical onset was recorded for every episode of clinical mastitis as well as for 8 other diseases defined using standardized case definitions (dystocia, milk fever, retained placenta, metritis, ketosis, displaced abomasum, lameness, and cystic ovarian disease) during the study period of October 1, 1999 to July 31, 2001. The mixed linear model for explaining variation in the outcome variable daily milk yield relative to non-mastitic herdmates found the terms for all 9 diseases studied, including clinical mastitis, significant. The model with an autoregressive correlation structure was preferred based on -2 * log likelihood, Akaike's information criterion, and Bayesian information criterion as well as savings in degrees of freedom. Separate analyses were run for first lactation cows and for second-plus lactation cows because their lactation curves were shaped differently. Adjusting for the effects of the other 8 diseases, milk production loss from clinical mastitis during the whole lactation was estimated as approximately 598 kg for second-plus lactation cows. However, cows that contracted mastitis had a daily production advantage of 2.6 kg over their herdmates until they contracted the disease. When compared with this potentially higher milk production, the total loss from clinical mastitis was estimated as 1181 kg.

Animals↗

Context-specific independence mixture modeling for positional weight matrices.

MOTIVATION: A positional weight matrix (PWM) is a statistical representation of the binding pattern of a transcription factor estimated from known binding site sequences. Previous studies showed that for factors which bind to divergent binding sites, mixtures of multiple PWMs increase performance. However, estimating a conventional mixture distribution for each position will in many cases cause overfitting. RESULTS: We propose a context-specific independence (CSI) mixture model and a learning algorithm based on a Bayesian approach. The CSI model adjusts complexity to fit the amount of variation observed on the sequence level in each position of a site. This not only yields a more parsimonious description of binding patterns, which improves parameter estimates, it also increases robustness as the model automatically adapts the number of components to fit the data. Evaluation of the CSI model on simulated data showed favorable results compared to conventional mixtures. We demonstrate its adaptive properties in a classical model selection setup. The increased parsimony of the CSI model was shown for the transcription factor Leu3 where two binding-energy subgroups were distinguished equally well as with a conventional mixture but requiring 30% less parameters. Analysis of the human-mouse conservation of predicted binding sites of 64 JASPAR TFs showed that CSI was as good or better than a conventional mixture for 89% of the TFs and for 70% for a single PWM model. AVAILABILITY: http://algorithmics.molgen.mpg.de/mixture.

Algorithms↗

Statistical approach to neural network model building for gentamicin peak predictions.

Feed forward neural networks are flexible, nonlinear modeling tools that are an extension of traditional statistical techniques. The hypothesis that feed forward neural network models can be built in a similar fashion as a statistical model was tested. Feed forward neural network models were built using forward and backward variable selection, and zero to five hidden nodes, and tanh and linear transfer functions were used. Gentamicin serum concentrations were predicted as a model drug for testing these methods. Peak observations from 392 patients were used to train, test, and validate the feed forward neural network. Inputs were demographic and drug dosing information. Model selection was performed using the Akaike information criteria (AIC), Bayesian information criteria (BIC), and a method of stopped training. The models with lowest root mean square (rms) error were those with all 10 inputs and five hidden nodes. Average rms error in the validation set was lowest for stopped training (1.46), then AIC (1.51), and finally BIC (1.56). Larger models tended to result in the best predictions. Overfitting can occur in models that are too large, either by using too many nodes in the hidden layer (rms = 1.49) or by using too many inputs with little information associated with them (rms = 1.70). We conclude that neural networks can be built using a large number of parameters that have good predictive performance. Care must be used during training to avoid overfitting the data. A stopped training method resulted in the network with the lowest rms error.

Adult↗

Bayesian analysis for the meiosis I non-disjunction fraction in numerical chromosomal anomalies.

The main causes of numerical chromosomal anomalies, including trisomies, arise from an error in the chromosomal segregation during the meiotic process, named a non-disjunction. One of the most used techniques to analyze chromosomal anomalies nowadays is the polymerase chain reaction (PCR), which counts the number of peaks or alleles in a polymorphic microsatellite locus. It was shown in previous works that the number of peaks has a multinomial distribution whose probabilities depend on the non-disjunction fraction F. In this work, we propose a Bayesian approach for estimating the meiosis I non-disjunction fraction F. in the absence of the parental information. Since samples of trisomic patients are, in general, small, the Bayesian approach can be a good alternative for solving this problem. We consider the sampling/importance resampling technique and the Simpson rule to extract information from the posterior distribution of F. Bayes and maximum likelihood estimators are compared through a Monte Carlo simulation, focusing on the influence of different sample sizes and prior specifications in the estimates. We apply the proposed method to estimate F. for patients with trisomy of chromosome 21 providing a sensitivity analysis for the method. The results obtained show that Bayes estimators are better in almost all situations.

Algorithms↗

Prediction of Nepsilon-acetylation on internal lysines implemented in Bayesian Discriminant Method.

Protein acetylation is an important and reversible post-translational modification (PTM), and it governs a variety of cellular dynamics and plasticity. Experimental identification of acetylation sites is labor-intensive and often limited by the availability of reagents such as acetyl-specific antibodies and optimization of enzymatic reactions. Computational analyses may facilitate the identification of potential acetylation sites and provide insights into further experimentation. In this manuscript, we present a novel protein acetylation prediction program named PAIL, prediction of acetylation on internal lysines, implemented in a BDM (Bayesian Discriminant Method) algorithm. The accuracies of PAIL are 85.13%, 87.97%, and 89.21% at low, medium, and high thresholds, respectively. Both Jack-Knife validation and n-fold cross-validation have been performed to show that PAIL is accurate and robust. Taken together, we propose that PAIL is a novel predictor for identification of protein acetylation sites and may serve as an important tool to study the function of protein acetylation. PAIL has been implemented in PHP and is freely available on a web server at: http://bioinformatics.lcd-ustc.org/pail.

Acetylation↗

Protein structure and fold prediction using Tree-Augmented naïve Bayesian classifier.

Due to the large volume of protein sequence data, computational methods to determine the structure class and the fold class of a protein sequence have become essential. Several techniques based on sequence similarity, Neural Networks, Support Vector Machines (SVMs), etc. have been applied. Since most of these classifiers use binary classifiers for multi-classification, there may be (N) c2 classifiers required. This paper presents a framework using the Tree-Augmented Bayesian Networks (TAN) which performs multi-classification based on the theory of learning Bayesian Networks and using improved feature vector representation of (Ding et al., 2001). In order to enhance TAN's performance, pre-processing of data is done by feature discretization and post-processing is done by using Mean Probability Voting (MPV) scheme. The advantage of using Bayesian approach over other learning methods is that the network structure is intuitive. In addition, one can read off the TAN structure probabilities to determine the significance of each feature (say, hydrophobicity) for each class, which helps to further understand the complexity in protein structure. The experiments on the datasets used in three prominent recent works show that our approach is more accurate than other discriminative methods. The framework is implemented on the BAYESPROT web server and it is available at http://www-appn.comp.nus.edu.sg/~bioinfo/bayesprot/Default.htm. More detailed results are also available on the above website.

Algorithms↗

The changing role of the exercise electrocardiogram as a diagnostic and prognostic test for chronic ischemic heart disease.

The exercise electrocardiogram has been the subject of intense research over the last 50 years, as both a diagnostic and prognostic method to assess patients with chronic ischemic heart disease. In 1986, the strengths and limitations of the technique to predict coronary and multivessel disease in clinical patient subsets are understood. The diagnostic accuracy of the test is improved by consideration of Bayesian theory, multivariate models and new non-ST segment criteria. Post-test coronary disease risk estimates are best reported in terms of a conditional probability, rather than statements of "positive" or "negative." The value of exercise testing in prognostic risk stratification is considerably enhanced by recent reports of long-term follow-up data in asymptomatic and symptomatic patients. Powerful prognostic information can be obtained when the clinical, electrocardiographic and physiologic data from the exercise test are used to formulate the post-test risk of a cardiac event, even in patients whose coronary anatomy is known. The changing role of the exercise electrocardiogram as a diagnostic and prognostic test is reviewed, with emphasis on the strengths and limitations of the procedure.

Angina Pectoris↗

Temporal relation between the ADC and DC potential responses to transient focal ischemia in the rat: a Markov chain Monte Carlo simulation analysis.

Markov chain Monte Carlo simulation was used in a reanalysis of the longitudinal data obtained by Harris et al. (J Cereb Blood Flow Metab 20:28-36) in a study of the direct current (DC) potential and apparent diffusion coefficient (ADC) responses to focal ischemia. The main purpose was to provide a formal analysis of the temporal relationship between the ADC and DC responses, to explore the possible involvement of a common latent (driving) process. A Bayesian nonlinear hierarchical random coefficients model was adopted. DC and ADC transition parameter posterior probability distributions were generated using three parallel Markov chains created using the Metropolis algorithm. Particular attention was paid to the within-subject differences between the DC and ADC time course characteristics. The results show that the DC response is biphasic, whereas the ADC exhibits monophasic behavior, and that the two DC components are each distinguishable from the ADC response in their time dependencies. The DC and ADC changes are not, therefore, driven by a common latent process. This work demonstrates a general analytical approach to the multivariate, longitudinal data-processing problem that commonly arises in stroke and other biomedical research.

Animals↗

Flat minima.

We present a new algorithm for finding low-complexity neural networks with high generalization capability. The algorithm searches for a "flat" minimum of the error function. A flat minimum is a large connected region in weight space where the error remains approximately constant. An MDL-based, Bayesian argument suggests that flat minima correspond to "simple" networks and low expected overfitting. The argument is based on a Gibbs algorithm variant and a novel way of splitting generalization error into underfitting and overfitting error. Unlike many previous approaches, ours does not require gaussian assumptions and does not depend on a "good" weight prior. Instead we have a prior over input-output functions, thus taking into account net architecture and training set. Although our algorithm requires the computation of second-order derivatives, it has backpropagation's order of complexity. Automatically, it effectively prunes units, weights, and input lines. Various experiments with feedforward and recurrent nets are described. In an application to stock market prediction, flat minimum search outperforms conventional backprop, weight decay, and "optimal brain surgeon/optimal brain damage".

Algorithms↗

Are grammatical representations useful for learning from biological sequence data?--a case study.

This paper investigates whether Chomsky-like grammar representations are useful for learning cost-effective, comprehensible predictors of members of biological sequence families. The Inductive Logic Programming (ILP) Bayesian approach to learning from positive examples is used to generate a grammar for recognising a class of proteins known as human neuropeptide precursors (NPPs). Collectively, five of the co-authors of this paper, have extensive expertise on NPPs and general bioinformatics methods. Their motivation for generating a NPP grammar was that none of the existing bioinformatics methods could provide sufficient cost-savings during the search for new NPPs. Prior to this project experienced specialists at SmithKline Beecham had tried for many months to hand-code such a grammar but without success. Our best predictor makes the search for novel NPPs more than 100 times more efficient than randomly selecting proteins for synthesis and testing them for biological activity. As far as these authors are aware, this is both the first biological grammar learnt using ILP and the first real-world scientific application of the ILP Bayesian approach to learning from positive examples. A group of features is derived from this grammar. Other groups of features of NPPs are derived using other learning strategies. Amalgams of these groups are formed. A recognition model is generated for each amalgam using C4.5 and C4.5rules and its performance is measured using both predictive accuracy and a new cost function, Relative Advantage (RA). The highest RA was achieved by a model which includes grammar-derived features. This RA is significantly higher than the best RA achieved without the use of the grammar-derived features. Predictive accuracy is not a good measure of performance for this domain because it does not discriminate well between NPP recognition models: despite covering varying numbers of (the rare) positives, all the models are awarded a similar (high) score by predictive accuracy because they all exclude most of the abundant negatives.

Bayes Theorem↗

Hidden Markov models for wavelet-based blind source separation.

In this paper, we consider the problem of blind source separation in the wavelet domain. We propose a Bayesian estimation framework for the problem where different models of the wavelet coefficients are considered: the independent Gaussian mixture model, the hidden Markov tree model, and the contextual hidden Markov field model. For each of the three models, we give expressions of the posterior laws and propose appropriate Markov chain Monte Carlo algorithms in order to perform unsupervised joint blind separation of the sources and estimation of the mixing matrix and hyper parameters of the problem. Indeed, in order to achieve an efficient joint separation and denoising procedures in the case of high noise level in the data, a slight modification of the exposed models is presented: the Bernoulli-Gaussian mixture model, which is equivalent to a hard thresholding rule in denoising problems. A number of simulations are presented in order to highlight the performances of the aforementioned approach: 1) in both high and low signal-to-noise ratios and 2) comparing the results with respect to the choice of the wavelet basis decomposition.

Algorithms↗

The Bayesian operating point of the Canny edge detector.

We have investigated the operating point of the Canny edge detector which minimizes the Bayes risk of misclassification. By considering each of the sequential stages which constitute the Canny algorithm, we conclude that the linear filtering stage of Canny, without postprocessing, performs very poorly by any standard in pattern recognition and achieves error rates which are almost indistinguishable from a priori classification. We demonstrate that the edge detection performance of the Canny detector is due almost entirely to the postprocessing stages of nonmaximal suppression and hysteresis thresholding.

Algorithms↗

Predicting gene regulation by sigma factors in Bacillus subtilis from genome-wide data.

MOTIVATION: Sigma factors regulate the expression of genes in Bacillus subtilis at the transcriptional level. We assess the accuracy of a fold-change analysis, Bayesian networks, dynamic models and supervised learning based on coregulation in predicting gene regulation by sigma factors from gene expression data. To improve the prediction accuracy, we combine sequence information with expression data by adding their log-likelihood scores and by using a logistic regression model. We use the resulting score function to discover currently unknown gene regulations by sigma factors. RESULTS: The coregulation-based supervised learning method gave the most accurate prediction of sigma factors from expression data. We found that the logistic regression model effectively combines expression data with sequence information. In a genome-wide search, highly significant logistic regression scores were found for several genes whose transcriptional regulation is currently unknown. We provide the corresponding RNA polymerase binding sites to enable a straightforward experimental verification of these predictions.

Algorithms↗

An MML classification of protein structure that knows about angles and sequence.

The MML classification program, Snob, deals with mixture modelling (or clustering) of circular data. It has recently been extended to do Markov modelling of the serial correlation between clusters such as modelling the fact that a Helix cluster favours being followed by another Helix cluster. Such a model is better known as a Hidden Markov Model. The search for the most appropriate secondary structure classification of protein data is of significant importance and was addressed by Hunter and States (1992) using the Bayesian classifier, AutoClass, on Cartesian co-ordinate data of protein residues. Dowe et al. (1996) improved upon this earlier work by using Snob to cluster dihedral angle data, with the advantage that 3 x 3 = 9 Cartesian co-ordinates can be represented by the 2 orientation-invariant angles, phi and psi. The Hidden Markov Model used here is shown to be a more appropriate way again of modelling protein data and results in the selection of a simpler class model with 17 structure classes. We report on this classification, including the class transition matrix, and relate it back to the amino-acid sequence and the simple Helix, Beta, Turn classification. We find 3 types of Helix, 2 types of Beta and many types of Turn. The msot numerous Turn class defines a continuous flexible structure that is negatively correlated to all the other classes.

Amino Acid Sequence↗

Variations over the message computation algorithm of lazy propagation.

Improving the performance of belief updating becomes increasingly important as real-world Bayesian networks continue to grow larger and more complex. In this paper, an investigation is done on how variations over the message-computation algorithm of lazy propagation may impact its performance. Lazy propagation is a junction-tree-based inference algorithm for belief updating in Bayesian networks. Lazy propagation combines variable elimination (VE) with a Shenoy-Shafer message-passing scheme in an attempt to exploit the independence properties induced by evidence in a junction-tree-based algorithm. The authors investigate, the use of arc reversal (AR) and symbolic probabilistic inference (SPI) as alternative algorithms for computing clique-to-clique messages in lazy propagation. The paper presents the results of an empirical evaluation of the performance of lazy propagation using AR, SPI, and VE as the message-computation algorithm. The results of the empirical evaluation show that no single algorithm outperforms or is outperformed by the other two alternatives. In many cases, there is no significant difference in the performance of the three algorithms.

Algorithms↗

The importance of proper model assumption in bayesian phylogenetics.

We studied the importance of proper model assumption in the context of Bayesian phylogenetics by examining >5,000 Bayesian analyses and six nested models of nucleotide substitution. Model misspecification can strongly bias bipartition posterior probability estimates. These biases were most pronounced when rate heterogeneity was ignored. The type of bias seen at a particular bipartition appeared to be strongly influenced by the lengths of the branches surrounding that bipartition. In the Felsenstein zone, posterior probability estimates of bipartitions were biased when the assumed model was underparameterized but were unbiased when the assumed model was overparameterized. For the inverse Felsenstein zone, however, both underparameterization and overparameterization led to biased bipartition posterior probabilities, although the bias caused by overparameterization was less pronounced and disappeared with increased sequence length. Model parameter estimates were also affected by model misspecification. Underparameterization caused a bias in some parameter estimates, such as branch lengths and the gamma shape parameter, whereas overparameterization caused a decrease in the precision of some parameter estimates. We caution researchers to assure that the most appropriate model is assumed by employing both a priori model choice methods and a posteriori model adequacy tests.

Bayes Theorem↗

Biomarker discovery in microarray gene expression data with Gaussian processes.

MOTIVATION: In clinical practice, pathological phenotypes are often labelled with ordinal scales rather than binary, e.g. the Gleason grading system for tumour cell differentiation. However, in the literature of microarray analysis, these ordinal labels have been rarely treated in a principled way. This paper describes a gene selection algorithm based on Gaussian processes to discover consistent gene expression patterns associated with ordinal clinical phenotypes. The technique of automatic relevance determination is applied to represent the significance level of the genes in a Bayesian inference framework. RESULTS: The usefulness of the proposed algorithm for ordinal labels is demonstrated by the gene expression signature associated with the Gleason score for prostate cancer data. Our results demonstrate how multi-gene markers that may be initially developed with a diagnostic or prognostic application in mind are also useful as an investigative tool to reveal associations between specific molecular and cellular events and features of tumour physiology. Our algorithm can also be applied to microarray data with binary labels with results comparable to other methods in the literature.

Algorithms↗