Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian computational modeling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

A flexible and fault tolerant query-reply system based on a Bayesian neural network.

A query-reply system based on a Bayesian neural network is described. Strategies for generating questions which make the system both efficient and highly fault tolerant are presented. This involves having one phase of question generation intended to quickly reach a hypothesis followed by a phase where verification of the hypothesis is attempted. In addition, both phases have strategies for detecting and removing inconsistencies in the replies from the user. Also described is an explanatory mechanism which gives information related to why a certain hypotheses is reached or question asked. Specific examples of the systems behavior as well as the results of a statistical evaluation are presented.

Animals↗

Linkage analysis of the simulated data - evaluations and comparisons of methods.

The goal of this study is to evaluate, compare, and contrast several standard and new linkage analysis methods. First, we compare a recently proposed confidence set approach with MAPMAKER/SIBS. Then, we evaluate a new Bayesian approach that accounts for heterogeneity. Finally, the newly developed software SIMPLE is compared with GENEHUNTER. We apply these methods to several replicates of the Genetic Analysis Workshop 13 simulated data to assess their ability to detect the high blood pressure genes on chromosome 21, whose positions were known to us prior to the analyses. In contrast to the standard methods, most of the new approaches are able to identify at least one of the disease genes in all the replicates considered.

Bayes Theorem↗

Increasing feasibility of optimal gene network estimation.

Disentangling networks of regulation of gene expression is a major challenge in the field of computational biology. Harvesting the information contained in microarray data sets is a promising approach towards this challenge. We propose an algorithm for the optimal estimation of Bayesian networks from microarray data, which reduces the CPU time and memory consumption of previous algorithms. We prove that the space complexity can be reduced from O(n(2) x 2(n)) to O(2(n)), and that the expected calculation time can be reduced from O(n(2) x 2(n)) to O(n x 2(n)), where n is the number of genes. We make intrinsic use of a limitation of the maximal number of regulators of each gene, which has biological as well as statistical justifications. The improvements are significant for some applications in research.

Algorithms↗

Bayesian model averaging: development of an improved multi-class, gene selection and classification tool for microarray data.

MOTIVATION: Selecting a small number of relevant genes for accurate classification of samples is essential for the development of diagnostic tests. We present the Bayesian model averaging (BMA) method for gene selection and classification of microarray data. Typical gene selection and classification procedures ignore model uncertainty and use a single set of relevant genes (model) to predict the class. BMA accounts for the uncertainty about the best set to choose by averaging over multiple models (sets of potentially overlapping relevant genes). RESULTS: We have shown that BMA selects smaller numbers of relevant genes (compared with other methods) and achieves a high prediction accuracy on three microarray datasets. Our BMA algorithm is applicable to microarray datasets with any number of classes, and outputs posterior probabilities for the selected genes and models. Our selected models typically consist of only a few genes. The combination of high accuracy, small numbers of genes and posterior probabilities for the predictions should make BMA a powerful tool for developing diagnostics from expression data. AVAILABILITY: The source codes and datasets used are available from our Supplementary website.

Algorithms↗

Comparison of methods for handling censored records in beef fertility data: field data.

The purpose of this study was to compare methods for handling censored days to calving records in beef cattle data, and verify results of an earlier simulation study. Data were records from natural service matings of 33,176 first-calf females in Australian Angus herds. Three methods for handling censored records were evaluated. Censored records (records on noncalving females) were assigned penalty values on a within-contemporary group basis under the first method (DCPEN). Under the second method (DCSIM), censored records were drawn from their respective predictive truncated normal distributions, whereas censored records were deleted under the third method (DCMISS). Data were analyzed using a mixed linear model that included the fixed effects of contemporary group and sex of calf, linear and quadratic covariates for age at mating, and random effects of animal and residual error. A Bayesian approach via Gibbs sampling was used to estimate variance components and predict breeding values. Posterior means (PM) (SD) of additive genetic variance for DCPEN, DCSIM, and DCMISS were 22.6d2 (4.2d2), 26.1d2 (3.6d2), and 13.5d2 (2.9d2), respectively. The PM (SD) of residual variance for DCPEN, DCSIM, and DCMISS were 431.4d2 (5.0d2), 371.4d2 (4.5d2), and 262.2d2 (3.4d2), respectively. The PM (SD) of heritability for DCPEN, DCSIM, and DCMISS were 0.05 (0.01), 0.07 (0.01), and 0.05 (0.01), respectively. Simulating trait records for noncalving females resulted in similar heritability to the penalty method but lower residual variance. Pearson correlations between posterior means of animal effects for sires with more than 20 daughters with records were 0.99 between DCPEN and DCSIM, 0.77 between DCPEN and DCMISS, and 0.81 between DCSIM and DCMISS. Of the 424 sires ranked in the top 10% and bottom 10% of sires in DCPEN, 91% and 89%, respectively, were also ranked in the top 10% and bottom 10% in DCSIM. Little difference was observed between DCPEN and DCSIM for correlations between posterior means of animal effects for sires, indicating that no major reranking of sires would be expected. This finding suggests little difference between these two censored data handling techniques for use in genetic evaluation of days to calving.

Animals↗

[An effective method for the estimation and comparison of the ED50 with small sample sizes].

In ED50 experiments the relationship between dose and probability of response is often modelled by the probit function. Standard statistical analysis estimates the parameters of this function by the maximum likelihood principle and derives the ED50 and its fiducial limits from these parameters. Bayesian analysis is more effective in two respects: It optionally includes prior information and in all but very few instances yields confidence intervals, whereas fiducial intervals often cannot be determined. Bayesian analysis of experiments with one substance has been treated in GRIEVE (1988). In the present article the mathematically interested reader is shown how to compare two substances. The probability of higher ED50 in the one substance as well as estimates of the ratio of the ED50's are obtained. The methods are easily extended to the effective dose for any other reasonable percentage of animals, e.g. ED90 or ED25. Experiments concerning lethal doses can be analysed by these methods as well. Both types of analysis are applied in two examples which compare new batches of vaccines with an established standard. In the first example both substances are nearly equivalent, while in the second example the new batch is considerably more efficient. An interactive FORTRAN program for a personal computer is available (cf. last section of 5.). It computes the maximum likelihood and the Bayesian solution, using approximate formulas in the latter case. Due to these approximations it was possible to develop a Bayesian program which is fast enough to run on a PC. Validation procedures have been performed. The output consists of a print file and, optionally, an ASCII file containing the coordinates of the posterior probability density and distribution functions.

Animals↗

Calcium regulation of single ryanodine receptor channel gating analyzed using HMM/MCMC statistical methods.

Type-II ryanodine receptor channels (RYRs) play a fundamental role in intracellular Ca(2+) dynamics in heart. The processes of activation, inactivation, and regulation of these channels have been the subject of intensive research and the focus of recent debates. Typically, approaches to understand these processes involve statistical analysis of single RYRs, involving signal restoration, model estimation, and selection. These tasks are usually performed by following rather phenomenological criteria that turn models into self-fulfilling prophecies. Here, a thorough statistical treatment is applied by modeling single RYRs using aggregated hidden Markov models. Inferences are made using Bayesian statistics and stochastic search methods known as Markov chain Monte Carlo. These methods allow extension of the temporal resolution of the analysis far beyond the limits of previous approaches and provide a direct measure of the uncertainties associated with every estimation step, together with a direct assessment of why and where a particular model fails. Analyses of single RYRs at several Ca(2+) concentrations are made by considering 16 models, some of them previously reported in the literature. Results clearly show that single RYRs have Ca(2+)-dependent gating modes. Moreover, our results demonstrate that single RYRs responding to a sudden change in Ca(2+) display adaptation kinetics. Interestingly, best ranked models predict microscopic reversibility when monovalent cations are used as the main permeating species. Finally, the extended bandwidth revealed the existence of novel fast buzz-mode at low Ca(2+) concentrations.

Algorithms↗

A genetic and spatial Bayesian analysis of mastitis resistance.

A nationwide health card recording system for dairy cattle was introduced in Norway in 1975 (the Norwegian Cattle Health Services). The data base holds information on mastitis occurrences on an individual cow basis. A reduction in mastitis frequency across the population is desired, and for this purpose risk factors are investigated. In this paper a Bayesian proportional hazards model is used for modelling the time to first veterinary treatment of clinical mastitis, including both genetic and environmental covariates. Sire effects were modelled as shared random components, and veterinary district was included as an environmental effect with prior spatial smoothing. A non-informative smoothing prior was assumed for the baseline hazard, and Markov chain Monte Carlo methods (MCMC) were used for inference. We propose a new measure of quality for sires, in terms of their posterior probability of being among the, say 10% best sires. The probability is an easily interpretable measure that can be directly used to rank sires. Estimating these complex probabilities is straightforward in an MCMC setting. The results indicate considerable differences between sires with regards to their daughters disease resistance. A regional effect was also discovered with the lowest risk of disease in the south-eastern parts of Norway.

Animals↗

Parallel Metropolis coupled Markov chain Monte Carlo for Bayesian phylogenetic inference.

MOTIVATION: Bayesian estimation of phylogeny is based on the posterior probability distribution of trees. Currently, the only numerical method that can effectively approximate posterior probabilities of trees is Markov chain Monte Carlo (MCMC). Standard implementations of MCMC can be prone to entrapment in local optima. Metropolis coupled MCMC [(MC)(3)], a variant of MCMC, allows multiple peaks in the landscape of trees to be more readily explored, but at the cost of increased execution time. RESULTS: This paper presents a parallel algorithm for (MC)(3). The proposed parallel algorithm retains the ability to explore multiple peaks in the posterior distribution of trees while maintaining a fast execution time. The algorithm has been implemented using two popular parallel programming models: message passing and shared memory. Performance results indicate nearly linear speed improvement in both programming models for small and large data sets.

Algorithms↗

[X-linked diseases and carrier detection].

Advances in molecular biology applied to the location of genes have generated a real evolution for the screening of women carrying X-related diseases. It is however imperative to first try to define the status of these women using classical methods: bayesian calculations taking into account genealogical data and direct screening which is partially reliable because of the inactivation of an X chromosome in women. The new genetic engineering techniques enable to locate the gene in an affected patient and follow its transmission in the families with the use of tracers linked to the gene of the disease. The difficulties of these studies are due to two phenomena. First, the risk of allele recombination because of a crossing over during the mitosis. This risk must be computed and specified. The second phenomenon is related to variations of information in the families which may either completely prevent identification of the carriers, or give less reliable results. The increasing number of molecular probes should enable to resolve this problem.

Bayes Theorem↗

Bayesian interval estimation of genetic relationships: application to paternity testing.

Using genetic marker data, we have developed a general methodology for estimating genetic relationships between a set of individuals. The purpose of this paper is to illustrate the practical utility of these methods as applied to the problem of paternity testing. Bayesian methods are used to compute the posterior probability distribution of the genetic relationship parameters. Use of an interval-estimation approach rather than a hypothesis-testing one avoids the problem of the specification of an appropriate null hypothesis in calculating the probability of paternity. Monte Carlo methods are used to evaluate the utility of two sets of genetic markers in obtaining suitably precise estimates of genetic relationship as well as the effect of the prior distribution chosen. Results indicate that with currently available markers a "true" father may be reliably distinguished from any other genetic relationship to the child and that with a reasonable number of markers one can often discriminate between an unrelated individual and one with a second-degree relationship to the child.

Bayes Theorem↗

Influence of biological variables upon pharmacokinetic parameters of intramuscular methotrexate in rheumatoid arthritis.

The pharmacokinetics of methotrexate were studied in 22 patients receiving 5-15 mg per week in a single i.m. administration for rheumatoid arthritis. The data consisted of 3 plasma levels per patient, taken at 2, 6, and 12 hours after the administration. The concentration of methotrexate was determined by fluorescence polarization immunoassay. The pharmacokinetic parameters of a 2-compartment model were determined by Bayesian estimation using the population values of Bressolle et al. [1996]. The fitted parameters were: total plasma clearance of methotrexate (CL), first-order absorption constant (ka), volume of central compartment (V1), and transfer constants between the 2 compartments (k12 and k21). Additional parameters were derived from the fitted ones: maximal concentration (Cmax), time to maximum (tmax), volume of distribution at steady-state (Vss), and terminal half-life (t1/2). Twenty-one biological covariates were considered to explain the interpatient variability. The relationships between these covariates and the pharmacokinetic parameters were investigated by principal component analysis and multiple regression analysis. About 90% of the variability of CL were explained by 4 variables (sex, age, height and serum creatinine). About 50%-70% of the variability of the other pharmacokinetic parameters were explained by a set of covariates including age, height, creatinine, creatinine clearance, and dose. The effect of dose was noticed mainly on k12, Vss, and t1/2, thus suggesting that the transfer of the drug from plasma to tissues may be nonlinear. The possibility of predicting CL with a good precision would facilitate the computation of dosage regimens in these patients.

Adult↗

Modularized learning of genetic interaction networks from biological annotations and mRNA expression data.

MOTIVATION: Inferring the genetic interaction mechanism using Bayesian networks has recently drawn increasing attention due to its well-established theoretical foundation and statistical robustness. However, the relative insufficiency of experiments with respect to the number of genes leads to many false positive inferences. RESULTS: We propose a novel method to infer genetic networks by alleviating the shortage of available mRNA expression data with prior knowledge. We call the proposed method 'modularized network learning' (MONET). Firstly, the proposed method divides a whole gene set to overlapped modules considering biological annotations and expression data together. Secondly, it infers a Bayesian network for each module, and integrates the learned subnetworks to a global network. An algorithm that measures a similarity between genes based on hierarchy, specificity and multiplicity of biological annotations is presented. The proposed method draws a global picture of inter-module relationships as well as a detailed look of intra-module interactions. We applied the proposed method to analyze Saccharomyces cerevisiae stress data, and found several hypotheses to suggest putative functions of unclassified genes. We also compared the proposed method with a whole-set-based approach and two expression-based clustering approaches.

Algorithms↗

Multilevel linear modelling for FMRI group analysis using Bayesian inference.

Functional magnetic resonance imaging studies often involve the acquisition of data from multiple sessions and/or multiple subjects. A hierarchical approach can be taken to modelling such data with a general linear model (GLM) at each level of the hierarchy introducing different random effects variance components. Inferring on these models is nontrivial with frequentist solutions being unavailable. A solution is to use a Bayesian framework. One important ingredient in this is the choice of prior on the variance components and top-level regression parameters. Due to the typically small numbers of sessions or subjects in neuroimaging, the choice of prior is critical. To alleviate this problem, we introduce to neuroimage modelling the approach of reference priors, which drives the choice of prior such that it is noninformative in an information-theoretic sense. We propose two inference techniques at the top level for multilevel hierarchies (a fast approach and a slower more accurate approach). We also demonstrate that we can infer on the top level of multilevel hierarchies by inferring on the levels of the hierarchy separately and passing summary statistics of a noncentral multivariate t distribution between them.

Bayes Theorem↗

A "holistic" kinesin phylogeny reveals new kinesin families and predicts protein functions.

Kinesin superfamily proteins are ubiquitous to all eukaryotes and essential for several key cellular processes. With the establishment of genome sequence data for a substantial number of eukaryotes, it is now possible for the first time to analyze the complete kinesin repertoires of a diversity of organisms from most eukaryotic kingdoms. Such a "holistic" approach using 486 kinesin-like sequences from 19 eukaryotes and analyzed by Bayesian techniques, identifies three new kinesin families, two new phylum-specific groups, and unites two previously identified families. The paralogue distribution suggests that the eukaryotic cenancestor possessed nearly all kinesin families. However, multiple losses in individual lineages mean that no family is ubiquitous to all organisms and that the present day distribution reflects common biology more than it does common ancestry. In particular, the distribution of four families--Kinesin-2, -9, and the proposed new families Kinesin-16 and -17--correlates with the possession of cilia/flagella, and this can be used to predict a flagellar function for two new kinesin families. Finally, we present a set of hidden Markov models that can reliably place most new kinesin sequences into families, even when from an organism at a great evolutionary distance from those in the analysis.

Animals↗

Independent origins of subgroup Bl + B2 and subgroup B3 metallo-beta-lactamases.

The metallo-beta-lactamases constitute Class B in the Ambler classification of beta-lactamases and are divided into three subclasses: Bl, B2, and B3. Bayesian phylogenies of the Subclass B1 + B2 and Subclass B3 metallo-beta-lactamases and their homologs show that the beta-lactam-hydrolyzing function evolved independently within each group. In Subclass B1+B2 that function evolved about 1 billion years ago, and in Subclass B3 it evolved before the divergence of the Gram-positive and Gram-negative eubacteria, about 2 billion years ago. These results lend additional support to the proposal that the metallo-beta-lactamases should be divided into two distinct classes.

Archaea↗

Tryptophanyl-tRNA synthetase crystal structure reveals an unexpected homology to tyrosyl-tRNA synthetase.

BACKGROUND: Tryptophanyl-tRNA synthetase (TrpRS) catalyzes activation of tryptophan by ATP and transfer to tRNA(Trp), ensuring translation of the genetic code for tryptophan. Interest focuses on mechanisms for specific recognition of both amino acid and tRNA substrates. RESULTS: Maximum-entropy methods enabled us to solve the TrpRS structure. Its three parts, a canonical dinucleotide-binding fold, a dimer interface, and a helical domain, have enough structural homology to tyrosyl-tRNA synthetase (TyrRS) that the two enzymes can be described as conformational isomers. Structure-based sequence alignment shows statistically significant genetic homology. Structural elements interacting with the activated amino acid, tryptophanyl-5'AMP, are almost exactly as seen in the TyrRS:tyrosyl-5'AMP complex. Unexpectedly, side chains that recognize indole are also highly conserved, and require reorientation of a 'specificity-determining' helix containing a conserved aspartate to assure selection of tryptophan versus tyrosine. The carboxy terminus, which is disordered and therefore not seen in TyrRS, forms part of the dimer interface in TrpRS. CONCLUSIONS: For the first time, the Bayesian statistical paradigm of entropy maximization and likelihood scoring has played a critical role in an X-ray structure solution. Sequence relatedness of structurally superimposable residues throughout TrpRS and TyrRS implies that they diverged more recently than most aminoacyl-tRNA synthetases. Subtle, tertiary structure changes are crucial for specific recognition of the two different amino acids. The conformational isomerism suggests that movement of the KMSKS loop, known to occur in the TyrRS transition state for amino acid activation, may provide a basis for conformational coupling during catalysis.

Amino Acid Sequence↗

Approximate likelihood-ratio test for branches: A fast, accurate, and powerful alternative.

We revisit statistical tests for branches of evolutionary trees reconstructed upon molecular data. A new, fast, approximate likelihood-ratio test (aLRT) for branches is presented here as a competitive alternative to nonparametric bootstrap and Bayesian estimation of branch support. The aLRT is based on the idea of the conventional LRT, with the null hypothesis corresponding to the assumption that the inferred branch has length 0. We show that the LRT statistic is asymptotically distributed as a maximum of three random variables drawn from the chi(0)2 + chi(1)2 distribution. The new aLRT of interior branch uses this distribution for significance testing, but the test statistic is approximated in a slightly conservative but practical way as 2(l1- l2), i.e., double the difference between the maximum log-likelihood values corresponding to the best tree and the second best topological arrangement around the branch of interest. Such a test is fast because the log-likelihood value l2 is computed by optimizing only over the branch of interest and the four adjacent branches, whereas other parameters are fixed at their optimal values corresponding to the best ML tree. The performance of the new test was studied on simulated 4-, 12-, and 100-taxon data sets with sequences of different lengths. The aLRT is shown to be accurate, powerful, and robust to certain violations of model assumptions. The aLRT is implemented within the algorithm used by the recent fast maximum likelihood tree estimation program PHYML (Guindon and Gascuel, 2003).

Classification↗