Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian modelling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Identifying drug active pathways from gene networks estimated by gene expression data.

We present a computational method for identifying genes and their regulatory pathways influenced by a drug, using microarray gene expression data collected by single gene disruptions and drug responses. The automatic identification of such genes and pathways in organisms' cells is an important problem for pharmacogenomics and the tailor-made medication. Our method estimates regulatory relationships between genes as a gene network from microarray data of gene disruptions with a Bayesian network model, then identifies the drug affected genes and their regulatory pathways on the estimated network with time course drug response microarray data. Compared to the existing method, our proposed method can identify not only the drug affected genes and the druggable genes, but also the drug responses of the pathways. For evaluating the proposed method, we conducted simulated examples based on artificial networks and expression data. Our method succeeded in identifying the pseudo drug affected genes and pathways with the high coverage greater than 80 %. We also applied our method to Saccharomyces cerevisiae drug response microarray data. In this real example, we identified the genes and the pathways that are potentially influenced by a drug. These computational experiments indicate that our method successfully identifies the drug-activated genes and pathways, and is capable of predicting undesirable side effects of the drug, identifying novel drug target genes, and understanding the unknown mechanisms of the drug.

Bayes Theorem↗

Analysis of the relationship between socioeconomic factors and stomach cancer incidence in Slovenia.

An unequal population distribution of well-known major risk factors explains much of the variation in the incidence of stomach cancer worldwide. The aim of this study was to determine whether geographical variation of the stomach cancer incidence rate between Slovenia's municipalities during years 1995-2001 could partially be explained by variations in the socioeconomic status as an indirect stomach cancer risk factor. A composite measure of each region's socioeconomic status, labelled as deprivation index, was created from basic socioeconomic characteristics of each municipality using factor analysis. Municipalities' standardized incidence ratios for all stomach cancers and non-cardia stomach cancer were calculated. A fully Bayesian spatial model with a conditionally autoregressive prior was applied using Markov chain Monte Carlo techniques and WinBUGS software. Spatially smoothed maps of stomach cancer incidence rates by 192 Slovenian municipalities show a clear west-to-east gradient. This pattern resembles the geographical variation of socioeconomic indices, but these indices are not significant predictors of stomach cancer incidence. Geographical variation of stomach cancer incidence in Slovenia could be partially explained by the heterogeneous socioeconomic characteristics of its municipalities. It is possible that the socioeconomic status indices used in our study were not enough powerful predictors of stomach cancer risk. Some further methodological research is needed to explain why this association was not statistically evident with the current modeling approach.

Adult↗

Gene networks as a tool to understand transcriptional regulation.

Gene regulatory networks, or simply gene networks (GNs), have shown to be a promising approach that the bioinformatics community has been developing for studying regulatory mechanisms in biological systems. GNs are built from the genome-wide high-throughput gene expression data that are often available from DNA microarray experiments. Conceptually, GNs are (un)directed graphs, where the nodes correspond to the genes and a link between a pair of genes denotes a regulatory interaction that occurs at transcriptional level. In the present study, we had two objectives: 1) to develop a framework for GN reconstruction based on a Bayesian network model that captures direct interactions between genes through nonparametric regression with B-splines, and 2) to demonstrate the potential of GNs in the analysis of expression data of a real biological system, the yeast pheromone response pathway. Our framework also included a number of search schemes to learn the network. We present an intuitive notion of GN theory as well as the detailed mathematical foundations of the model. A comprehensive analysis of the consistency of the model when tested with biological data was done through the analysis of the GNs inferred for the yeast pheromone pathway. Our results agree fairly well with what was expected based on the literature, and we developed some hypotheses about this system. Using this analysis, we intended to provide a guide on how GNs can be effectively used to study transcriptional regulation. We also discussed the limitations of GNs and the future direction of network analysis for genomic data. The software is available upon request.

Bayes Theorem↗

Using Bayesian networks to predict survival of liver transplant patients.

The relative scarcity of grafts available for liver transplantation highlights the need to identify patients likely to have good outcomes after treatment. We used transplant information from the United Network for Organ Sharing database to construct a Bayesian network model to predict 90-day graft survival. The final model incorporated a set of 29 pre-transplant variables, and it achieved performance, as measured by area under the receiver operating characteristic curve, of 0.674 by cross-validation and 0.681 on an independent validation set. The results showed a positive predictive value of 91%, while the negative predictive value was lower at 30%. With additional refinement and validation, our model may be useful as an adjunct to clinical experience in identifying patients most likely to have good outcomes following liver transplantation.

Adult↗

Urban farming and malaria risk factors in a medium-sized town in Cote d'Ivoire.

Urbanization occurs at a rapid pace across Africa and Asia and affects people's health and well-being. A typical feature in urban settings of Africa is the maintenance of traditional livelihoods, including agriculture. The purpose of this study was to investigate malaria risk factors in urban farming communities in a medium-sized town in Côte d'Ivoire. Two cross-sectional surveys were carried out among 112 households from six agricultural zones. First, the heads of households were interviewed on agricultural land use, farming practices, water storage, sanitation facilities, and socioeconomic status. Second, a finger prick blood sample was taken from all household members and examined for the occurrence and density of Plasmodia. Geographic coordinates of houses, farming plots, and potential mosquito breeding sites were recorded and integrated into a geographic information system. Predictors of Plasmodium falciparum parasitemia were assessed using non-random and random effects Bayesian regression models. The overall prevalence of P. falciparum was 32.1%. In children < 15 years of age, risk factors for a P. falciparum infection included living in a specific agricultural zone, close proximity to permanent ponds and fish ponds, periodic stays overnight in temporary farm huts, and low socioeconomic status. Our findings indicate that specific crop systems and specific agricultural practices may increase the risk of malaria in urban settings of tropical Africa.

Agriculture↗

Transplantation statistics in the UK--an agenda for the next quinquennium.

Our next quinquennial plan for transplantation studies extends MPI to corneal and unrelated marrow transplantation. It applies Bayesian hierarchical modelling to regional variation in donor procurement and continues a program of special studies to augment national databases with respect to kidney, corneal, heart and liver transplantation. It promotes research collaboration among European and other organ exchange organizations as pioneered in the Council of Europe 1986 Study on High Sensitization, which showed the effectiveness of the network of European organ exchange organizations in liaising with transplant units.

Bayes Theorem↗

Suicide risk prediction by computer interview: a prospective study.

A computer interview program that uses a subjective Bayesian probability model to assess suicide risk was evaluated. Predictions made by clinicians for 52 patients were compared with predictions made by the computer for the same patients. The computer was significantly (p = .001) better at predicting attempters, and clinicians were significantly (p = .01) better at predicting nonattempters. An analysis of receiver operating characteristic curves showed that the computer had better overall discrimination, but the difference was nonsignificant.

Decision Making, Computer-Assisted↗

Clinical inferences and decisions--II. Decision trees, receiver operator curves and subjective probability.

In patient management, clinical decisions follow a logical sequence which can be formally expressed as a decision tree in which the uncertainties associated with each alternative outcome may be made explicit using Bayes' theorem. Where test data is used in the formulation of a decision, the uncertainty associated with the information it conveys may be modified by changing the pass/fail criterion to alter the false positive and false negative error rate. Classical procedures based on information theory are described to illustrate how this may be achieved for any test. When hard data is not available to permit such an approach, the clinician must rely on his own past experience or that of a colleague. Several methods are available for quantifying such experience by estimating subjective probabilities associated with an action or test result. Two simple methods are described for deriving subjective probabilities for subsequent use within a Bayesian decision model.

Bayes Theorem↗

Case-control diagnosis and Bayesian inference in common viral infections.

The predictive values of symptoms and signs for given diseases are often unknown. The fact that a high proportion of individuals with a certain disease may have a specific group of symptoms (the case-control approach) does not necessarily mean that the specific group of symptoms will allow one reliably to diagnose the disease. This study, utilizing a population based data set for common acute infections, shows that descriptions of common viral illnesses found in medical textbooks that associate illnesses with symptoms do not allow one to predict reliably isolation of the supposed causal organism. Positive predictive value of groups of symptoms for specific viral infections did not exceed 11 percent in this study. However, the data closely fitted the Bayesian statistical model often proposed for such decision making by physicians.

Adolescent↗

[The prognostic evaluation of acute pancreatitis].

Early identification of severity is one of the most important problems in acute pancreatitis, both for decision-making and classification. Predictive criteria show a wide range of accuracy: clinical examination (on admission: 76-85%); single laboratory data (PCR: 68-98%, C3-C4: 63-72%); multifactorial scoring systems (Ranson: 65-82%, Imrie: 78-95%); diagnostic peritoneal lavage (72-90%); CT features (52-81%). In 1982 we started a prospective evaluation of the prognostic performances of a bayesian statistical model for the prediction of severe vs mild pancreatis and death vs survival, which uses the outcome-related patterns of several variables, assuming their independence, analysed on a data of 44 patients. The performances have been calculated prospectively by comparing the expected vs actual results on 88 further patients (accuracy, sensitivity and specificity, respectively, in the prediction of severe pancreatitis: 92%, 92%, 93%; in the prediction of death: 95%, 97%, 87%). Moreover, the model can represent classes of risk by combining prediction of death + severe pancreatitis (DSP), survival + severe pancreatitis (SSP) and survival + mild pancreatitis (SMP) (accuracy, sensitivity and specificity, respectively, in the prediction of DSP: 97%, 83%, 100%; in the prediction of SSP: 95%, 87%, 97%; in the prediction of SMP: 95%, 97%, 90%). Our model enables clinicians dealing with other population to re-determine different variables or integrate them with new information, whenever available. It seems to be transferable and adaptable, even with a probable further increase of the performances, without compromising the objectivity of the predictive judgement and the homogeneity of the classes of risk.

Acute Disease↗

Large hierarchical Bayesian analysis of multivariate survival data.

Failure times that are grouped according to shared environments arise commonly in statistical practice. That is, multiple responses may be observed for each of many units. For instance, the units might be patients or centers in a clinical trial setting. Bayesian hierarchical models are appropriate for data analysis in this context. At the first stage of the model, survival times can be modelled via the Cox partial likelihood, using a justification due to Kalbfleisch (1978, Journal of the Royal Statistical Society, Series B 40, 214-221). Thus, questionable parametric assumptions are avoided. Conventional wisdom dictates that it is comparatively safe to make parametric assumptions at subsequent stages. Thus, unit-specific parameters are modelled parametrically. The posterior distribution of parameters given observed data is examined using Markov chain Monte Carlo methods. Specifically, the hybrid Monte Carlo method, as described by Neal (1993a, in Advances in Neural Information Processing 5, 475-482; 1993b, Probabilistic inference using Markov chain Monte Carlo methods), is utilized.

Antineoplastic Combined Chemotherapy Protocols↗

Advances in meta-analysis as a research method.

Meta-analysis is a method of data analysis applied to summarizing research findings, both quantitative and qualitative summaries of individual studies. Rooted in the fundamental values of the scientific enterprise of replicability, causal and correlational analysis, it is useful for answering three general questions: What is the central tendency or typical study outcome? How much variability exists among study outcomes? What is the explanation of the variability? Advanced statistical and mathematical techniques are being used in counting studies with significant and non-significant findings; in combining effect size estimates based on fixed- or random-effects models; in the use of general linear models; and, Bayesian procedures. The methodology of meta-analysis and suggestions for publications are presented as well as appropriate software programs available for use in the meta-analytic process are explored. The benefits of meta-analysis as a research method in effecting health policy development is offered as a pragmatic perspective for future consideration. Conclusions have implications for the use of meta-analysis as a teaching strategy and as a methodology in nursing research and other applied sciences. In view of the rapid pace of knowledge development and increased public demand for accountability, meta-analysis offers an opportunity for organizing phenomena which gives direction for provision of quality health care.

Humans↗

Change-point analysis of neuron spike train data.

In many medical experiments, data are collected across time, over a number of similar trials, or over a number of experimental units. As is the case of neuron spike train studies, these data may be in the form of counts of events per unit of time. These counts may be correlated within each trial. It is often of interest to know if the introduction of an intervention, such as the application of a stimulus, affects the distribution of the counts over the course of the experiment. In such investigations, each trial generates a sequence of data that may or may not contain a change in distribution at some point in time. Each sequence of integer counts can be viewed as arising from a Poisson process and are therefore independently distributed or as an integer-valued time series that allows for correlations between these counts. The main aim of this paper is to show how the ensemble of sample paths may be used to make inference about the distribution of the instantaneous times of change in a given population. This will be accomplished using a Bayesian hierarchical model for these change-points in time. A bonus of these models is they also allow for inference about the probability of a change in each unit and the magnitude of the effects, if any. The use of such change-point models on integer-valued time series is illustrated on neuron spike train data, although the methods can be applied to other situations where integer-valued processes arise.

Action Potentials↗

Bayesian methods for a three-state model for rodent carcinogenicity studies.

The objective of a chronic rodent bioassay is to assess the impact of a chemical compound on the development of tumors. However, most tumor types are not observable prior to necropsy, making direct estimation of the tumor incidence rate problematic. In such cases, estimation can proceed only if the study incorporates multiple interim sacrifices or we make use of simplified parametric or nonparametric models. In addition, it is widely accepted that other factors, such as weight, can be related to both dose level and tumor onset, confounding the association of interest. However, there is not typically enough information in the current study to assess such effects. The addition of historical data can help alleviate this problem. In this article, we propose a novel Bayesian semiparametric model for the analysis of data from rodent carcinogenicity studies. We develop informative prior distributions for covariate effects through the use of historical control data and outline a Gibbs sampling scheme. We implement the model by analyzing data from a National Toxicology Program chronic rodent bioassay.

Animals↗

Bayesian identification of differentially expressed isoforms using a novel joint model of RNA-seq data.

We develop a Bayesian approach, BayesIso, to identify differentially expressed isoforms from RNA-seq data. The approach features a novel joint model of the sample variability and the deferential state of isoforms. Specifically, the within-sample variability and the between-sample variability of each isoform are modeled by a Poisson-Lognormal model and a Gamma-Gamma model, respectively. Using a Bayesian framework, the differential state of each isoform and the model parameters are jointly estimated by a Markov Chain Monte Carlo (MCMC) method. Extensive studies using simulation and real data demonstrate that BayesIso can effectively detect isoforms of less differentially expressed and differential transcripts for genes with multiple isoforms. We applied the approach to breast cancer RNA-seq data and uncovered a unique set of isoforms that form key pathways associated with breast cancer recurrence. First, PI3K/AKT/mTOR signaling and PTEN signaling pathways are identified as being involved in breast cancer development. Further integrated with protein-protein interaction data, pathways of Jak-STAT, mTOR, MAPK and Wnt signaling are revealed in association with breast cancer recurrence. Finally, several pathways are activated in the early recurrence of breast cancer. In tumors that occur early, members of pathways of cellular metabolism and cell cycle (such as CD36 and TOP2A) are upregulated, while immune response genes such as NFATC1 are downregulated.

Humans↗

Multilevel linear modelling for FMRI group analysis using Bayesian inference.

Functional magnetic resonance imaging studies often involve the acquisition of data from multiple sessions and/or multiple subjects. A hierarchical approach can be taken to modelling such data with a general linear model (GLM) at each level of the hierarchy introducing different random effects variance components. Inferring on these models is nontrivial with frequentist solutions being unavailable. A solution is to use a Bayesian framework. One important ingredient in this is the choice of prior on the variance components and top-level regression parameters. Due to the typically small numbers of sessions or subjects in neuroimaging, the choice of prior is critical. To alleviate this problem, we introduce to neuroimage modelling the approach of reference priors, which drives the choice of prior such that it is noninformative in an information-theoretic sense. We propose two inference techniques at the top level for multilevel hierarchies (a fast approach and a slower more accurate approach). We also demonstrate that we can infer on the top level of multilevel hierarchies by inferring on the levels of the hierarchy separately and passing summary statistics of a noncentral multivariate t distribution between them.

Bayes Theorem↗

Comparing artificial and convolutional neural networks with traditional models for Genomic prediction in wheat.

With the rapid development of sequencing technology, the application of genomic prediction has become more and more common in breeding schemes of livestocks and crops. Selecting an appropriate statistical model is of central importance to achieve high prediction accuracy. Recently, machine learning models have been expected to upgrade genomic prediction into a new era. However, the perspective still suffers from lack of evidence that machine learning models can generally outperform the traditional ones on empirical data sets. In this study, we compared two machine learning models based on artificial neural network (ANN) and convolutional neural network (CNN) with four traditional models, including genomic best linear unbiased prediction (GBLUP), Bayesian ridge regression (BRR), BayesA and BayesB, using three published data sets for grain yield in wheat. For each model, we considered two variants: modeling and ignoring the genotype-by-environment ([Formula: see text]) interaction. In the comparison, we considered two strategies of cross-validation: predicting genotypes that have not been evaluated in any environment (CV1) and predicting genotypes that have been tested in other environments (CV2). Our results showed that traditional Bayesian models (BayesA, BayesB, and BRR) outperformed GBLUP, ANN and CNN when considering [Formula: see text] interaction. The accuracies of ANN and CNN were higher than traditional models only in CV1 and when [Formula: see text] interaction was ignored. It was also found that the performance of the two machine learning models was significantly affected by the interaction between the CV strategy and the way of treating the [Formula: see text] interaction, while that of the four traditional models was only influenced by whether the [Formula: see text] interaction was considered or not. Thus, machine learning models can be a powerful complementary to the traditional ones and their superiority may depend on the prediction scenario. Among the two machine learning models, we observed that the accuracy of ANN was higher than CNN in most cases, indicating that it is still challenging to adapt complex machine learning models such as CNN to genomic prediction.

ANN↗