Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Semiparametric regression modeling with mixtures of Berkson and classical error, with application to fallout from the Nevada test site.

We construct Bayesian methods for semiparametric modeling of a monotonic regression function when the predictors are measured with classical error. Berkson error, or a mixture of the two. Such methods require a distribution for the unobserved (latent) predictor, a distribution we also model semiparametrically. Such combinations of semiparametric methods for the dose response as well as the latent variable distribution have not been considered in the measurement error literature for any form of measurement error. In addition, our methods represent a new approach to those problems where the measurement error combines Berkson and classical components. While the methods are general, we develop them around a specific application, namely, the study of thyroid disease in relation to radiation fallout from the Nevada test site. We use this data to illustrate our methods, which suggest a point estimate (posterior mean) of relative risk at high doses nearly double that of previous analyses but that also suggest much greater uncertainty in the relative risk.

Bayes Theorem↗

The RIN: an RNA integrity number for assigning integrity values to RNA measurements.

BACKGROUND: The integrity of RNA molecules is of paramount importance for experiments that try to reflect the snapshot of gene expression at the moment of RNA extraction. Until recently, there has been no reliable standard for estimating the integrity of RNA samples and the ratio of 28S:18S ribosomal RNA, the common measure for this purpose, has been shown to be inconsistent. The advent of microcapillary electrophoretic RNA separation provides the basis for an automated high-throughput approach, in order to estimate the integrity of RNA samples in an unambiguous way. METHODS: A method is introduced that automatically selects features from signal measurements and constructs regression models based on a Bayesian learning technique. Feature spaces of different dimensionality are compared in the Bayesian framework, which allows selecting a final feature combination corresponding to models with high posterior probability. RESULTS: This approach is applied to a large collection of electrophoretic RNA measurements recorded with an Agilent 2100 bioanalyzer to extract an algorithm that describes RNA integrity. The resulting algorithm is a user-independent, automated and reliable procedure for standardization of RNA quality control that allows the calculation of an RNA integrity number (RIN). CONCLUSION: Our results show the importance of taking characteristics of several regions of the recorded electropherogram into account in order to get a robust and reliable prediction of RNA integrity, especially if compared to traditional methods.

Algorithms↗

Influence of network topology and data collection on network inference.

We recently developed an approach for testing the accuracy of network inference algorithms by applying them to biologically realistic simulations with known network topology. Here, we seek to determine the degree to which the network topology and data sampling regime influence the ability of our Bayesian network inference algorithm, NETWORKINFERENCE, to recover gene regulatory networks. NETWORKINFERENCE performed well at recovering feedback loops and multiple targets of a regulator with small amounts of data, but required more data to recover multiple regulators of a gene. When collecting the same number of data samples at different intervals from the system, the best recovery was produced by sampling intervals long enough such that sampling covered propagation of regulation through the network but not so long such that intervals missed internal dynamics. These results further elucidate the possibilities and limitations of network inference based on biological data.

Algorithms↗

Probabilistic independent component analysis for functional magnetic resonance imaging.

We present an integrated approach to probabilistic independent component analysis (ICA) for functional MRI (FMRI) data that allows for nonsquare mixing in the presence of Gaussian noise. In order to avoid overfitting, we employ objective estimation of the amount of Gaussian noise through Bayesian analysis of the true dimensionality of the data, i.e., the number of activation and non-Gaussian noise sources. This enables us to carry out probabilistic modeling and achieves an asymptotically unique decomposition of the data. It reduces problems of interpretation, as each final independent component is now much more likely to be due to only one physical or physiological process. We also describe other improvements to standard ICA, such as temporal prewhitening and variance normalization of timeseries, the latter being particularly useful in the context of dimensionality reduction when weak activation is present. We discuss the use of prior information about the spatiotemporal nature of the source processes, and an alternative-hypothesis testing approach for inference, using Gaussian mixture models. The performance of our approach is illustrated and evaluated on real and artificial FMRI data, and compared to the spatio-temporal accuracy of results obtained from classical ICA and GLM analyses.

Algorithms↗

Preoperative prediction of malignancy of ovarian tumors using least squares support vector machines.

In this work, we develop and evaluate several least squares support vector machine (LS-SVM) classifiers within the Bayesian evidence framework, in order to preoperatively predict malignancy of ovarian tumors. The analysis includes exploratory data analysis, optimal input variable selection, parameter estimation, and performance evaluation via receiver operating characteristic (ROC) curve analysis. LS-SVM models with linear and radial basis function (RBF) kernels, and logistic regression models have been built on 265 training data, and tested on 160 newly collected patient data. The LS-SVM model with nonlinear RBF kernel achieves the best performance, on the test set with the area under the ROC curve (AUC), sensitivity and specificity equal to 0.92, 81.5% and 84.0%, respectively. The best averaged performance over 30 runs of randomized cross-validation is also obtained by an LS-SVM RBF model, with AUC, sensitivity and specificity equal to 0.94, 90.0% and 80.6%, respectively. These results show that the LS-SVM models have the potential to obtain a reliable preoperative distinction between benign and malignant ovarian tumors, and to assist the clinicians for making a correct diagnosis.

Adnexal Diseases↗

Interval estimation of urban ozone level and selection of influential factors by employing automatic relevance determination model.

In this work, we focus on simulating the ground-level ozone (O3) time series and its daily maximum concentration in Hong Kong urban air by employing the multilayer perceptron (MLP) model combined with the automatic relevance determination (ARD) method (for simplicity, we name it as MLP-ARD model). Two air quality monitoring sites in Hong Kong, i.e., Tsuen Wan and Tung Chung, are selected for the numerical experiments. The MLP-ARD model based on Bayesian evidence framework can provide reliable interval estimation of real observation as well as offering efficient strategy to avoid over-fitting. The performance comparisons between MLP-ARD model and traditional artificial neural network (ANN) model based on maximum likelihood indicate that MLP-ARD model is more powerful to capture the wild fluctuation of O3 level especially during O3 episodes than the traditional model. Furthermore, it can assess and rank the input variables for the prediction according to their relative importance to the output variable, i.e., the daily maximum O3 concentration in this study. The preliminary experimental results indicate that nitric oxide (NO) and solar radiation are the most important input variables for O3 prediction at both selected sites. In addition, the previous daily maximum O3 level is also important for Tung Chung site. In this regard, MLP-ARD model is a feasible tool to interpret the real physical and chemical process of urban O3 variation.

Air Pollutants↗

Prediction of protein function using protein-protein interaction data.

Assigning functions to novel proteins is one of the most important problems in the postgenomic era. Several approaches have been applied to this problem, including the analysis of gene expression patterns, phylogenetic profiles, protein fusions, and protein-protein interactions. In this paper, we develop a novel approach that employs the theory of Markov random fields to infer a protein's functions using protein-protein interaction data and the functional annotations of protein's interaction partners. For each function of interest and protein, we predict the probability that the protein has such function using Bayesian approaches. Unlike other available approaches for protein annotation in which a protein has or does not have a function of interest, we give a probability for having the function. This probability indicates how confident we are about the prediction. We employ our method to predict protein functions based on "biochemical function," "subcellular location," and "cellular role" for yeast proteins defined in the Yeast Proteome Database (YPD, www.incyte.com), using the protein-protein interaction data from the Munich Information Center for Protein Sequences (MIPS, mips.gsf.de). We show that our approach outperforms other available methods for function prediction based on protein interaction data. The supplementary data is available at www-hto.usc.edu/~msms/ProteinFunction.

Bayes Theorem↗

Accuracy of cDNA microarray methods to detect small gene expression changes induced by neuregulin on breast epithelial cells.

BACKGROUND: cDNA microarrays are a powerful means to screen for biologically relevant gene expression changes, but are often limited by their ability to detect small changes accurately due to "noise" from random and systematic errors. While experimental designs and statistical analysis methods have been proposed to reduce these errors, few studies have tested their accuracy and ability to identify small, but biologically important, changes. Here, we have compared two cDNA microarray experimental design methods with northern blot confirmation to reveal changes in gene expression that could contribute to the early antiproliferative effects of neuregulin on MCF10AT human breast epithelial cells. RESULTS: We performed parallel experiments on identical samples using a dye-swap design with ANOVA and an experimental design that excludes systematic biases by "correcting" experimental/control hybridization ratios with control/control hybridizations on a spot-by-spot basis. We refer to this approach as the "control correction method" (CCM). Using replicate arrays, we identified a decrease in proliferation genes and an increase in differentiation genes. Using an arbitrary cut-off of 1.7-fold and p values <0.05, we identified a total of 32 differentially expressed genes, 9 with the dye-swap method, 18 with the CCM, and 5 genes with both methods. 23 of these 32 genes were subsequently verified by northern blotting. Most of these were <2-fold changes. While the dye-swap method (using either ANOVA or Bayesian analysis) detected a smaller number of genes (14-16) compared to the CCM (46), it was more accurate (89-92% vs. 75%). Compared to the northern blot results, for most genes, the microarray results underestimated the fold change, implicating the importance of detecting these small changes. CONCLUSIONS: We validated two experimental design paradigms for cDNA microarray experiments capable of detecting small (<2-fold) changes in gene expression with excellent fidelity that revealed potentially important genes associated with the anti-proliferative effects of neuregulin on MCF10AT breast epithelial cells.

Blotting, Northern↗

Applying dynamic Bayesian networks to perturbed gene expression data.

BACKGROUND: A central goal of molecular biology is to understand the regulatory mechanisms of gene transcription and protein synthesis. Because of their solid basis in statistics, allowing to deal with the stochastic aspects of gene expressions and noisy measurements in a natural way, Bayesian networks appear attractive in the field of inferring gene interactions structure from microarray experiments data. However, the basic formalism has some disadvantages, e.g. it is sometimes hard to distinguish between the origin and the target of an interaction. Two kinds of microarray experiments yield data particularly rich in information regarding the direction of interactions: time series and perturbation experiments. In order to correctly handle them, the basic formalism must be modified. For example, dynamic Bayesian networks (DBN) apply to time series microarray data. To our knowledge the DBN technique has not been applied in the context of perturbation experiments. RESULTS: We extend the framework of dynamic Bayesian networks in order to incorporate perturbations. Moreover, an exact algorithm for inferring an optimal network is proposed and a discretization method specialized for time series data from perturbation experiments is introduced. We apply our procedure to realistic simulations data. The results are compared with those obtained by standard DBN learning techniques. Moreover, the advantages of using exact learning algorithm instead of heuristic methods are analyzed. CONCLUSION: We show that the quality of inferred networks dramatically improves when using data from perturbation experiments. We also conclude that the exact algorithm should be used when it is possible, i.e. when considered set of genes is small enough.

Algorithms↗

Face verification across age progression.

Human faces undergo considerable amounts of varialions with aging. While face recognition systems have been proven to be sensitive to factors such as illumination and pose, their sensitivity to facial aging effects is yet to be studied. How does age progression affect the similarity between a pair of face images of an individual? What is the confidence associated with establishing the identity between a pair of age separated face images? In this paper, we develop a Bayesian age difference classifier that classifies face images of individuals based on age differences and performs face verification across age progression. Further, we study the similarity of faces across age progression. Since age separated face images invariably differ in illumination and pose, we propose preprocessing methods for minimizing such variations. Experimental results using a database comprising of pairs of face images that were retrieved from the passports of 465 individuals are presented. The verification system for faces separated by as many as nine years, attains an equal error rate of 8.5%.

Adult↗

Language-tree divergence times support the Anatolian theory of Indo-European origin.

Languages, like genes, provide vital clues about human history. The origin of the Indo-European language family is "the most intensively studied, yet still most recalcitrant, problem of historical linguistics". Numerous genetic studies of Indo-European origins have also produced inconclusive results. Here we analyse linguistic data using computational methods derived from evolutionary biology. We test two theories of Indo-European origin: the 'Kurgan expansion' and the 'Anatolian farming' hypotheses. The Kurgan theory centres on possible archaeological evidence for an expansion into Europe and the Near East by Kurgan horsemen beginning in the sixth millennium BP. In contrast, the Anatolian theory claims that Indo-European languages expanded with the spread of agriculture from Anatolia around 8,000-9,500 years bp. In striking agreement with the Anatolian hypothesis, our analysis of a matrix of 87 languages with 2,449 lexical items produced an estimated age range for the initial Indo-European divergence of between 7,800 and 9,800 years bp. These results were robust to changes in coding procedures, calibration points, rooting of the trees and priors in the bayesian analysis.

Agriculture↗

Genrate: a generative model that finds and scores new genes and exons in genomic microarray data.

Recently, researchers have made some progress in using microarrays to validate predicted exons in genome sequence and find new gene structures. However, current methods rely on separately making threshold-based decisions on intensity of expression, similarity of expression profiles, and arrangements of exons in the genome. We have taken a Bayesian approach and developed GenRate, a generative model that accounts for both genome-wide expression data taken from multiple conditions (e.g. tissues) and co-location and density of probes in DNA sequence data. GenRate balances probabilistic evidence derived from different sources and outputs scores (log-likelihoods) for each gene model, enabling the estimation of false-positive and false-negative rates. The model has a number of local minima that is exponential in the length of the DNA sequence data, so direct application of the EM learning algorithm produces poor results. We describe a novel way of parameterizing the model using examples from the data set, so that good solutions are found using an efficient algorithm. We apply GenRate to a subset of mouse genome-wide expression data that we have created, and discuss the statistical significance of the genes found by GenRate. Three of the highest-ranking gene structures found by GenRate, each containing thousands of bases from the genome, are confirmed using RT-PCR experiments.

Animals↗

A graphic nomogram for warfarin dosage adjustment.

We assessed the ability of a graphic nomogram to adjust steady-state warfarin dosages and to predict international normalized ratios (INR) after a dosage change, compared with an anticoagulation clinic pharmacist and a Bayesian regression computer program. Study subjects were 108 men and 3 women receiving warfarin anticoagulation. In all patients the median absolute errors in predicted INR values for the nomogram, computer program, and pharmacist were 0.33, 0.46, and 0.48, respectively. The nomogram was significantly more precise than both other methods (p=0.036). In a subset of 50 patients who required dosage reductions, the median absolute INR prediction errors for the nomogram, computer program, and pharmacist were 0.35, 0.54, and 0.48 respectively. The nomogram was significantly more precise than the pharmacist (p=0.005) and computer (p=0.002). The ability to provide more precise dosage reductions of warfarin may be of clinical importance in light of current recommendations for higher-intensity warfarin therapy and maintenance of higher INR values. Prospective validation of the performance of this nomogram in a routine clinical setting is warranted.

Adult↗

Evaluation of machine-learning methods for ligand-based virtual screening.

Machine-learning methods can be used for virtual screening by analysing the structural characteristics of molecules of known (in)activity, and we here discuss the use of kernel discrimination and naive Bayesian classifier (NBC) methods for this purpose. We report a kernel method that allows the processing of molecules represented by binary, integer and real-valued descriptors, and show that it is little different in screening performance from a previously described kernel that had been developed specifically for the analysis of binary fingerprint representations of molecular structure. We then evaluate the performance of an NBC when the training-set contains only a very few active molecules. In such cases, a simpler approach based on group fusion would appear to provide superior screening performance, especially when structurally heterogeneous datasets are to be processed.

Artificial Intelligence↗

Bayesian analysis of mixtures applied to post-synaptic potential fluctuations.

Bayesian inference techniques have been applied to the analysis of fluctuation of post-synaptic potentials in the hippocampus. The underlying statistical model assumes that the varying synaptic signals are characterized by mixtures of (unknown) numbers of individual gaussian, or normal, component distributions. Each solution consists of a group of individual components with unique mean values and relative probabilities of occurrence and a predictive probability density. The advantages of bayesian inference techniques over the alternative method of maximum likelihood estimation (MLE) of the parameters of an unknown mixture distribution include the following: (1) prior information may be incorporated in the estimation of model parameters; (2) conditional probability estimates of the number of individual components in the mixture are calculated; (3) flexibility exists in the extent to which the estimated noise standard deviation indicates the width of each component; (4) posterior distributions for component means are calculated, including measures of uncertainty about the means; and (5) probability density functions of the component distributions and the overall mixture distribution are estimated in relation to the raw grouped data, together with measures of uncertainty about these estimates. This expository report describes this novel approach to the unconstrained identification of components within a mixture, and provides demonstration of the usefulness of the technique in the context of both simulations and the analysis of distributions of synaptic potential signals.

Action Potentials↗

Spatiotemporal motion boundary detection and motion boundary velocity estimation for tracking moving objects with a moving camera: a level sets PDEs approach with concurrent camera motion compensation.

The purpose of this study is to investigate a method of tracking moving objects with a moving camera. This method estimates simultaneously the motion induced by camera movement. The problem is formulated as a Bayesian motion-based partitioning problem in the spatiotemporal domain of the image quence. An energy functional is derived from the Bayesian formulation. The Euler-Lagrange descent equations determine imultaneously an estimate of the image motion field induced by camera motion and an estimate of the spatiotemporal motion undary surface. The Euler-Lagrange equation corresponding to the surface is expressed as a level-set partial differential equation for topology independence and numerically stable implementation. The method can be initialized simply and can track multiple objects with nonsimultaneous motions. Velocities on motion boundaries can be estimated from geometrical properties of the motion boundary. Several examples of experimental verification are given using synthetic and real-image sequences.

Algorithms↗

Performance-based selection of likelihood models for phylogeny estimation.

Phylogenetic estimation has largely come to rely on explicitly model-based methods. This approach requires that a model be chosen and that that choice be justified. To date, justification has largely been accomplished through use of likelihood-ratio tests (LRTs) to assess the relative fit of a nested series of reversible models. While this approach certainly represents an important advance over arbitrary model selection, the best fit of a series of models may not always provide the most reliable phylogenetic estimates for finite real data sets, where all available models are surely incorrect. Here, we develop a novel approach to model selection, which is based on the Bayesian information criterion, but incorporates relative branch-length error as a performance measure in a decision theory (DT) framework. This DT method includes a penalty for overfitting, is applicable prior to running extensive analyses, and simultaneously compares all models being considered and thus does not rely on a series of pairwise comparisons of models to traverse model space. We evaluate this method by examining four real data sets and by using those data sets to define simulation conditions. In the real data sets, the DT method selects the same or simpler models than conventional LRTs. In order to lend generality to the simulations, codon-based models (with parameters estimated from the real data sets) were used to generate simulated data sets, which are therefore more complex than any of the models we evaluate. On average, the DT method selects models that are simpler than those chosen by conventional LRTs. Nevertheless, these simpler models provide estimates of branch lengths that are more accurate both in terms of relative error and absolute error than those derived using the more complex (yet still wrong) models chosen by conventional LRTs. This method is available in a program called DT-ModSel.

Bayes Theorem↗

Image registration using a symmetric prior--in three dimensions.

This paper describes a Bayesian method for three-dimensional registration of brain images. A finite element approach is used to obtain a maximum a posteriori estimate of the deformation field at every voxel of a template volume. The priors used by the MAP estimate penalize unlikely deformations and enforce a continuous one-to-one mapping. The deformations are assumed to have some form of symmetry, in that priors describing the probability distribution of the deformations should be identical to those for the inverses (i.e., warping brain A to brain B should not be different probablistically from warping B to A). A gradient descent algorithm is presented for estimating the optimum deformations.

Algorithms↗