Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

A Bayesian threshold-normal mixture model for analysis of a continuous mastitis-related trait.

Mastitis is associated with elevated somatic cell count in milk, inducing a positive correlation between milk somatic cell score (SCS) and the absence or presence of the disease. In most countries, selection against mastitis has focused on selecting parents with genetic evaluations that have low SCS. Univariate or multivariate mixed linear models have been used for statistical description of SCS. However, an observation of SCS can be regarded as drawn from a 2- (or more) component mixture defined by the (usually) unknown health status of a cow at the test-day on which SCS is recorded. A hierarchical 2-component mixture model was developed, assuming that the health status affecting the recorded test-day SCS is completely specified by an underlying liability variable. Based on the observed SCS, inferences can be drawn about disease status and parameters of both SCS and liability to mastitis. The prior probability of putative mastitis was allowed to vary between subgroups (e.g., herds, families), by specifying fixed and random effects affecting both SCS and liability. Using simulation, it was found that a Bayesian model fitted to the data yielded parameter estimates close to their true values. The model provides selection criteria that are more appealing than selection for lower SCS. The proposed model can be extended to handle a wide range of problems related to genetic analyses of mixture traits.

Animals↗

Inferring the structure of latent class models using a genetic algorithm.

Present optimization techniques in latent class analysis apply the expectation maximization algorithm or the Newton-Raphson algorithm for optimizing the parameter values of a prespecified model. These techniques can be used to find maximum likelihood estimates of the parameters, given the specified structure of the model, which is defined by the number of classes and, possibly, fixation and equality constraints. The model structure is usually chosen on theoretical grounds. A large variety of structurally different latent class models can be compared using goodness-of-fit indices of the chi-square family, Akaike's information criterion, the Bayesian information criterion, and various other statistics. However, finding the optimal structure for a given goodness-of-fit index often requires a lengthy search in which all kinds of model structures are tested. Moreover, solutions may depend on the choice of initial values for the parameters. This article presents a new method by which one can simultaneously infer the model structure from the data and optimize the parameter values. The method consists of a genetic algorithm in which any goodness-of-fit index can be used as a fitness criterion. In a number of test cases in which data sets from the literature were used, it is shown that this method provides models that fit equally well as or better than the models suggested in the original articles.

Algorithms↗

Modelling developmental instability as the joint action of noise and stability: a Bayesian approach.

BACKGROUND: Fluctuating asymmetry is assumed to measure individual and population level developmental stability. The latter may in turn show an association with stress, which can be observed through asymmetry-stress correlations. However, the recent literature does not support an ubiquitous relationship. Very little is known why some studies show relatively strong associations while others completely fail to find such a correlation. We propose a new Bayesian statistical framework to examine these associations RESULTS: We are considering developmental stability - i.e. the individual buffering capacity - as the biologically relevant trait and show that (i) little variation in developmental stability can explain observed variation in fluctuating asymmetry when the distribution of developmental stability is highly skewed, and (ii) that a previously developed tool (i.e. the hypothetical repeatability of fluctuating asymmetry) contains only limited information about variation in developmental stability, which stands in sharp contrast to the earlier established close association between the repeatability and developmental instability. CONCLUSION: We provide tools to generate valuable information about the distribution of between-individual variation in developmental stability. A simple linear transformation of a previous model lead to completely different conclusions. Thus, theoretical modelling of asymmetry and stability appears to be very sensitive to the scale of inference. More research is urgently needed to get better insights in the developmental mechanisms of noise and stability. In spite of the fact that the model is likely to represent an oversimplification of reality, the accumulation of new insights could be incorporated in the Bayesian statistical approach to obtain more reliable estimation.

Bayes Theorem↗

Inferring the population structure and demography of Drosophila ananassae from multilocus data.

Inferring the origin, population structure, and demographic history of a species is a major objective of population genetics. Although many organisms have been analyzed, the genetic structures of subdivided populations are not well understood. Here we analyze Drosophila ananassae, a highly substructured, cosmopolitan, and human-commensal species distributed in the tropical, subtropical, and mildly temperate regions of the world. We adopt a multilocus approach (with 10 neutral loci) using 16 population samples covering almost the entire species range (Asia, Australia, and America). Analyzed with our recently developed Bayesian method, 5 populations in Southeast Asia are found to be central, while the other 11 are peripheral. These 5 central populations were sampled from localities that belonged to a single landmass ("Sundaland") during the late Pleistocene ( approximately 18,000 years ago), when sea level was approximately 120 m below the present level. The inferred migration routes of D. ananassae out of Sundaland seem to parallel those of humans in this region. Strong evidence for a population size expansion is seen particularly in the ancestral populations.

Animals↗

Linking phylogenetics with population genetics to reconstruct the geographic origin of a species.

Reconstructing ancestral geographic origins is critical for understanding the long-term evolution of a species. Bayesian methods have been proposed to test biogeographic hypotheses while accommodating uncertainty in phylogenetic reconstruction. However, the problem that certain taxa may have a disproportionate influence on conclusions has not been addressed. Here, we infer the geographic origin of Drosophila simulans using 2,014 bp of the period locus from 63 lines collected from 18 countries. We also analyze two previously published datasets, alcohol dehydrogenase related and NADH:ubiquinone reductase 75 kDa subunit precursor. Phylogenetic inferences of all three loci support Madagascar as the geographic origin of D. simulans. Our phylogenetic conclusions are robust to taxon resampling and to the potentially confounding effects of recombination. To test our phylogenetically derived hypothesis we develop a randomization test of the population genetics prediction that sequences from the geographic origin should contain more genetic polymorphism than those from derived populations. We find that the Madagascar population has elevated genetic polymorphism relative to non-Madagascar sequences. These data are corroborated by mitochondrial DNA sequence data.

Animals↗

Linked vs unlinked markers: multilocus microsatellite haplotype-sharing as a tool to estimate gene flow and introgression.

We have explored the use of multilocus microsatellite haplotypes to study introgression from cultivated (Malus domestica) into wild apple (Malus sylvestris), and to study gene flow among remnant populations of M. sylvestris. A haplotype consisted of alleles at microsatellite loci along one chromosome. As destruction of haplotypes through recombination occurs much faster than loss of alleles due to genetic drift, the lifespan of a multilocus haplotype is much shorter than that of the underlying alleles. When different populations share the same haplotype, this may indicate recent gene flow between populations. Similarly, haplotypes shared between two species would be a strong signal for introgression. As the expected lifespan of a haplotype depends on the strength of the linkage, the length [in centiMorgans (cM)] of the haplotype shared contains information on the number of generations passed. This application of shared haplotypes is distinct from using haplotype-sharing to detect association between markers and a certain trait. We inferred haplotypes for four to eight microsatellite loci on Linkage Group 10 of apple from genotype data using the program phase, and then identified those haplotypes shared between populations and species. Compared with a Bayesian analysis of unlinked microsatellite loci using the program structure, haplotype-sharing detected a partially different set of putative hybrids. Cultivated haplotypes present in M. sylvestris were short (< 1.5 cM), indicating that introgression had taken place many generations ago, except for two Belgian plants that contained a haplotype of 47.1 cM, indicating recent introgression. In the estimation of gene flow, F(ST) based on unlinked loci indicated small (0.032-0.058) but statistically significant differentiation between some populations only. However, various M. sylvestris haplotypes were shared in nearly all pairwise comparisons of populations, and their length indicated recent gene flow. Hence, all Dutch populations should be considered as one conservation unit. The added value of using sharing of multilocus microsatellite haplotypes as a source of population genetic information is discussed.

Belgium↗

Comparison of methods for handling censored records in beef fertility data: simulation study.

A simulation study was conducted to compare methods for handling censored records for days to calving in beef cattle data. Days to calving was defined as the time, in days, between when a bull is turned out in the pasture and the subsequent parturition. Simulated data were generated to have data structure and genetic relationships similar to an available field data set. Records were simulated for 33,176 daughters of 4,238 sires. Data were simulated using a mixed linear model that included the fixed effects of contemporary group and sex of calf, linear and quadratic covariates for age at mating, and random effects of animal and residual error. Two methods for handling censored records were evaluated, and two censoring rates of 12 and 20% were applied to assess the influence of higher censoring rates on inferences. Censored records were assigned penalty values on a within-contemporary group basis under the first method (DCPEN). Under the second method (DCSIM), censored records were drawn from their respective predictive distributions. A Bayesian approach via Gibbs sampling was used to estimate variance components and predict breeding values. Posterior means (PM) and standard deviations (SD) of additive genetic variance for DCPEN at 12 and 20% censoring were 23.2 (3.7) and 21.0 (3.6), respectively, whereas the same estimates for DCSIM at 12 and 20% censoring were 23.7(3.3) and 21.9 (3.4), respectively. In all cases, the true value of the genetic variance was within the high posterior density (HPD) interval (95%). The PM (SD) of residual variance for DCPEN at 12 and 20% censoring were 415.7 (4.7) and 440.0 (4.8) respectively, whereas the same estimates for DCSIM at 12 and 20% censoring were 371.0 (4.3) and 365.4 (4.4), respectively. The true value of the residual variance was within the HPD (95%) for DCSIM, but it was outside this interval for DCPEN at both censoring rates, indicating a systematic bias for this parameter. Bayes Factor and Deviance Information Criteria were used for model comparisons, and both criteria indicated the superiority of the DCSIM method. However, little difference was observed between the two methods for correlations between true breeding values and posterior means of animal effects for sires, indicating that no major reranking of sires would be expected. This finding suggests that either censored data handling technique can be successfully used in a genetic evaluation for days to calving.

Animals↗

Mapping a quantitative trait locus via the EM algorithm and Bayesian classification.

Mapping a locus controlling a quantitative genetic trait (e.g., blood pressure) to a specific genomic region is of considerable interest. Data on the quantitative trait under consideration and several codominant genetic markers with known genomic locations are collected from members of families and statistically analyzed to draw inferences on the genomic position of the trait locus. The vector of parameters of interest comprises the pairwise recombination fractions, theta, between the putative quantitative trait locus and the marker loci. One of the major complications in estimating theta for a quantitative trait in humans is the lack of haplotype information on members of families. The purpose of this study was to devise a computationally simple and efficient method of estimation of theta in the absence of haplotype information. We have proposed a two-stage estimation procedure using the expectation-maximization (EM) algorithm. In the first stage, parameters of the QTL are estimated based on data of a sample of unrelated individuals. From estimates thus obtained, we have used a Bayes' rule to infer QTL genotypes of parents in families. Finally, in the second stage of the procedure, we have proposed an EM algorithm for obtaining the maximum likelihood estimate of theta based on data of informative families (which are identified upon inferring parental QTL genotypes performed in the first stage). We have shown, using simulated data, that the proposed procedure is cost-effective, computationally simple, and statistically efficient. As expected, analysis of data on multiple markers jointly is more efficient than the analysis based on single markers.

Algorithms↗

Evolutionary timescale of rabies virus adaptation to North American bats inferred from the substitution rate of the nucleoprotein gene.

Throughout North America, rabies virus (RV) is endemic in bats. Distinct RV variants exist that are closely associated with infection of individual host species, such that there is little or no sustained spillover infection away from the primary host. Using Bayesian methodology, nucleotide substitution rates were estimated from alignments of partial nucleoprotein (N) gene sequences of nine distinct bat RV variants from North America. Substitution rates ranged from 2.32 x 10(-4) to 1.38 x 10(-3) substitutions per site per year. A maximum-likelihood (ML) molecular clock model was rejected for only two of the nine datasets. In addition, using sequences from bat RV variants across the Americas, the evolutionary rate for the complete N gene was estimated to be 2.32 x 10(-4). This rate was used to scale trees using Bayesian and ML methods, and the time of the most recent common ancestor for current bat RV variant diversity in the Americas was estimated to be 1660 (range 1267-1782) and 1651 (range 1254-1773), respectively. Our reconstructions suggest that RV variants currently associated with infection of bats from Latin America (Desmodus and Tadarida) share the earliest common ancestor with the progenitor RV. In addition, from the ML tree, times were estimated for the emergence of the three major lineages responsible for bat rabies cases in North America. Adaptation to infection of the colonial bat species analysed (Eptesicus fuscus, Myotis spp.) appears to have occurred much quicker than for the solitary species analysed (Lasionycteris noctivagans, Pipistrellus subflavus, Lasiurus borealis, Lasiurus cinereus), suggesting that the process of virus adaptation may be dependent on host biology.

Adaptation, Biological↗

Catarrhine primate divergence dates estimated from complete mitochondrial genomes: concordance with fossil and nuclear DNA evidence.

Accurate divergence date estimates improve scenarios of primate evolutionary history and aid in interpretation of the natural history of disease-causing agents. While molecule-based estimates of divergence dates of taxa within the superfamily Hominoidea (apes and humans) are common in the literature, few such estimates are available for the Cercopithecoidea (Old World monkeys), the sister taxon of the hominoids in the primate infraorder Catarrhini. To help fill this gap, we have sequenced the entire mitochondrial DNA (mtDNA) genomes from a representative of three cercopithecoid tribes, Cercopithecini (Chlorocebus aethiops), Colobini (Colobus guereza), and Presbytini (Trachypithecus obscurus), and analyzed these new data together with other catarrhine mtDNA genomes available in public databases. Molecular divergence date estimates are dependent on calibration points gleaned from the paleontological record. We defined criteria for the selection of good calibration points and identified three points meeting these criteria: Homo-Pan, 6.0 Ma; Pongo-hominines, 14.0 Ma; hominoid/cercopithecoid, 23.0 Ma. Because a uniform molecular clock does not fit the catarrhine mtDNA data, we estimated divergence dates using a penalized likelihood and a Bayesian method, both of which take into account the effects of rate differences on lineages, phylogenetic tree structure, and multiple calibration points. The penalized likelihood method applied to the coding regions of the mtDNA genome yielded the following divergence date estimates, with approximate 95% confidence intervals: cercopithecine-colobine, 16.2 (14.4-17.9) Ma; colobin-presbytin, 10.9 (9.6-12.3) Ma; cercopithecin-papionin, 11.6 (10.3-12.9) Ma; and Macaca-Papio, 9.8 (8.6-10.9) Ma. Within the hominoids, the following dates were inferred: hylobatid-hominid, 16.8 (15.0-18.5) Ma; Gorilla-Homo+Pan, 8.1 (7.1-9.0) Ma; Pongo pygmaeus pygmaeus-P. p. abelii, 4.1 (3.5-4.7) Ma; and Pan troglodytes-P. paniscus, 2.4 (2.0-2.7) Ma. These dates were similar to those found using penalized likelihood on other subsets of the data, but slightly younger than several of the Bayesian estimates.

Africa↗

Nonlinear local electrovascular coupling. II: From data to neuronal masses.

In the companion article a local electrovascular coupling (LEVC) model was proposed to explain the continuous dynamics of electrical and vascular states within a cortical unit. These states produce certain mesoscopic reflections whose discrete time series can be reconstructed from electroencephalography (EEG) and functional magnetic resonance imaging (fMRI). In this article we develop a recursive optimization algorithm based on the local linearization (LL) filter and an innovation method to make statistical inferences about the LEVC model from both EEG and fMRI data, i.e., to estimate the unobserved states and the unknown parameters of the model. For a better understanding, the LL filter is described from a Bayesian point of view, providing the particulars for the case of hybrid data (e.g., EEG and fMRI), which could be sampled at different rates. The dynamics of the exogenous synaptic inputs going into the cortical unit are also estimated by introducing a set of Gaussian radial basis functions. In order to study the dynamics of the electrical and vascular states in the striate cortex of humans as well as their local interrelationships, we applied this algorithm to EEG and fMRI recordings obtained concurrently from two subjects while passively observing a radial checkerboard with a white/black pattern reversal. The EEG and fMRI data from the first subject was used to estimate the electrical/vascular states and parameters of the LEVC model in V1 for a 4.0 Hz reversion frequency. We used the EEG data from the second subject to investigate the changes in the dynamics of the electrical states when the frequency of reversion is varied from 0.5-4.0 Hz. Then we made use of the estimated electrical states to predict the effects on the vasculature that such variations produce.

Bayes Theorem↗

Connective molecular pathways of experimental bladder inflammation.

Inflammation is an inherent response of the organism that permits its survival despite constant environmental challenges. The process normally leads to recovery from injury and to healing. However, if targeted destruction and assisted repair are not properly phased, chronic inflammation can result in persistent tissue damage. To better understand the inflammatory process, we recently introduced a profiling methodology to identify common genes involved in bladder inflammation. The method represents a complementation to the classic quantification of inflammation and provides information regarding the early, intermediate, and late events in gene regulation. However, gene profiling fails to describe the molecular pathways and their interconnections involved in the particular inflammatory response. The present work introduces a new statistical technique for inferring functional interconnections between inflammatory pathways underlying classic models of bladder inflammation and permits the modeling of the inflammatory network. This new statistical method is based on variants of cluster analysis, Boolean networking, differential equations, Bayesian networking, and partial correlation. By applying partial correlation analysis, we developed mosaics of gene expression that permitted a global visualization of common and unique pathways elicited by different stimuli. The significance of these processes was tested from both biological and statistical viewpoints. We propose that connective mosaic may represent the necessary simplification step to visualize cDNA array results.

Animals↗

Interim analyses in clinical research.

Interim analyses are those that occur before the scheduled completion of a clinical study. The motivation for such analyses may be to see whether conclusive evidence is available concerning the aims of the study, or it may be simple curiosity. Statisticians disagree about the impact that interim analyses have on inferences that can be drawn from the study. The significance testing view insists that conclusions from a study with interim analyses are different than from one without--even though the data are identical. In the Bayesian view there is no penalty for interim analyses: study results can even be monitored continually without changing the conclusions. Both views are explained and recommendations for designing and analyzing studies with interim analyses are made.

Bayes Theorem↗

Mitochondrial data support an odd-nosed colobine clade.

To obtain a more complete understanding of the evolutionary history of the leaf-eating monkeys we have examined the mitochondrial genome sequence of two African and six Asian colobines. Although taxonomists have proposed grouping the "odd-nosed" colobines (proboscis monkey, douc langur, and the snub-nosed monkey) together, phylogenetic support for such a clade has not been tested using molecular data. Phylogenetic analyses using parsimony, maximum likelihood, and Bayesian methods support a monophyletic clade of odd-nosed colobines consisting of Nasalis, Pygathrix, and Rhinopithecus, with tentative support for Nasalis occupying a basal position within this clade. The African and Asian colobine lineages are inferred to have diverged by 10.8 million years ago (mya or Ma). Within the Asian colobines the odd-nosed clade began to diversify by 6.7 Ma. These results augment our understanding of colobine evolution, particularly the nature and timing of the colobine expansion into Asia. This phylogenetic information will aid those developing conservation strategies for these highly endangered, diverse, and unique primates.

Animals↗

Determination of an optimal dosage regimen using a Bayesian decision analysis of efficacy and adverse effect data.

One of the aims of Phase II clinical trials is to determine the dosage regimen(s) that will be investigated during a confirmatory Phase III clinical trial. During Phase II, pharmacodynamic data are collected that enables the efficacy and safety of the drug to be assessed. It is proposed in this paper to use Bayesian decision analysis to determine the optimal dosage regimen based on efficacy and toxicity of the drug oxybutynin used in the treatment of urinary urge incontinence. Such an approach results in a general framework allowing modeling, inference and decision making to be carried out. For oxybutynin, the repeated measurement efficacy and toxicity data were modeled using nonlinear hierarchical models and inferences were based on posterior probabilities. The optimal decision in this problem was to determine the dosage regimen that maximized the posterior expected utility given the prior information on the model parameters and the patient response data. The utility function was defined using clinical opinion on the satisfactory levels of efficacy and toxicity and then combined by weighting the relative importance of each pharmacodynamic response. Markov chain Monte Carlo (MCMC) methodology implemented in Win-BUGS 1.3 was used to obtain posterior estimates of the model parameters, probabilities and utilities.

Aged↗

Skeletal growth estimation using radiographic image processing and analysis.

An automated knowledge-based vision system for skeletal growth estimation in children is reported in this paper. Images were obtained from hand radiographs of 32 male and 25 female children of age 1-16 yr. Phalanx bones were automatically localized and segmented using hierarchical inferences and active shape models, respectively. A number of shape descriptors were obtained from the segmented bone contour to quantify skeletal growth. From these descriptors, a feature vector was selected for a regression model and a Bayesian estimator. The estimation accuracy was 84% for females and 82% for males. This level of accuracy is comparable to that of expert pediatric radiologists, which suggests that the proposed approach has a potential application in pediatric medicine.

Adolescent↗

A Bayesian approach to DNA sequence segmentation.

Many deoxyribonucleic acid (DNA) sequences display compositional heterogeneity in the form of segments of similar structure. This article describes a Bayesian method that identifies such segments by using a Markov chain governed by a hidden Markov model. Markov chain Monte Carlo (MCMC) techniques are employed to compute all posterior quantities of interest and, in particular, allow inferences to be made regarding the number of segment types and the order of Markov dependence in the DNA sequence. The method is applied to the segmentation of the bacteriophage lambda genome, a common benchmark sequence used for the comparison of statistical segmentation algorithms.

Algorithms↗

A neural network approach to approximating MAP in belief networks.

Bayesian belief networks (BBN) are a widely studied graphical model for representing uncertainty and probabilistic interdependence among variables. One of the factors that restricts the model's wide acceptance in practical applications is that the general inference with BBN is NP-hard. This is also true for the maximum a posteriori probability (MAP) problem, which is to find the most probable joint value assignment to all uninstantiated variables, given instantiation of some variables in a BBN. To circumvent the difficulty caused by MAP's computational complexity, we suggest in this paper a neural network approximation approach. With this approach, a BBN is treated as a neural network without any change or transformation of the network structure, and the node activation functions are derived based on an energy function defined over a given BBN. Three methods are developed. They are the hill-climbing style discrete method, the simulated annealing method, and the continuous method based on the mean field theory. All three methods are for BBN of general structures, with the restriction that nodes of BBN are binary variables. In addition, rules for applying these methods to noisy-or networks are also developed, which may lead to more efficient computation in some cases. These methods' convergence is analyzed, and their validity tested through a series of computer experiments with two BBN of moderate size and complexity. Although additional theoretical and empirical work is needed, the analysis and experiments suggest that this approach may lead to effective and accurate approximation for MAP problems.

Algorithms↗