Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Phylogenetic relationships among anchovies, sardines, herrings and their relatives (Clupeiformes), inferred from whole mitogenome sequences.

The relationships among and within the main lineages of the order Clupeiformes have been explored in few morphological studies and still remain poorly understood. Using whole mitogenome sequences, we inferred the relationships among 25 clupeiform species, sampled from each clupeiform family and subfamily, and a large selection of non-clupeiform teleosts. Our character sets, including unambiguously aligned, concatenated mitogenome sequences that we have divided into four (1st and 2nd codon positions, tRNA genes, and rRNA genes) or five partitions (same as before plus the transversions at 3rd codon positions, using 'RY' coding), were analyzed by the partitioned Bayesian method. The result strongly supported the monophyly of the Clupeiformes within the Otocephala, with Denticeps clupeoides as the sister group of a clade comprising all the remaining clupeiforms species (= suborder Clupeoidei). Within the Clupeoidei, the family Engraulidae was the sister group of the remaining taxa, comprising members of Sundasalangidae, Pristigasteridae, Clupeidae and Chirocentridae. Relationships among the latter four families remained ambiguous. In particular, the position of the Chirocentridae was difficult to estimate possibly owing to its higher molecular evolutionary rate. Of the five subfamilies in the family Clupeidae, monophylies of three (Alosinae, Clupeinae and Dorosomatinae) were statistically rejected. Instead, our mitogenomic data provide strong support for new clades within the Clupeidae, some of which are composed of members of more than one of the previously accepted subfamilies.

Animals↗

Advances in statistical methods to map quantitative trait loci in outbred populations.

Statistical methods to map quantitative trait loci (QTL) in outbred populations are reviewed, extensions and applications to human and plant genetic data are indicated, and areas for further research are identified. Simple and computationally inexpensive methods include (multiple) linear regression of phenotype on marker genotypes and regression of squared phenotypic differences among relative pairs on estimated proportions of identity-by-descent at a locus. These methods are less suited for genetic parameter estimation in outbred populations but allow the determination of test statistic distributions via simulation or data permutation; however, further inferences including confidence intervals of QTL location require the use of Monte Carlo or bootstrap sampling techniques. A method which is intermediate in computational requirements is residual maximum likelihood (REML) with a covariance matrix of random QTL effects conditional on information from multiple linked markers. Testing for the number of QTLs on a chromosome is difficult in a classical framework. The computationally most demanding methods are maximum likelihood and Bayesian analysis, which take account of the distribution of multilocus marker-QTL genotypes on a pedigree and permit investigators to fit different models of variation at the QTL. The Bayesian analysis includes the number of QTLs on a chromosome as an unknown.

Bayes Theorem↗

Localization of a quantitative trait locus via a Bayesian approach.

A Bayesian approach to the direct mapping of a quantitative trait locus (QTL), fully utilizing information from multiple linked gene markers, is presented in this paper. The joint posterior distribution (a mixture distribution modeling the linkage between a biallelic QTL and N gene markers) is computationally challenging and invites exploration via Markov chain Monte Carlo methods. The parameter's complete marginal posterior densities are obtained, allowing a diverse range of inferences. Parameters estimated include the QTL genotype probabilities for the sires and the offspring, the allele frequencies for the QTL, and the position and additive and dominance effects of the QTL. The methodology is applied through simulation to a half-sib design to form an outbred pedigree structure where there is an entire class of missing information. The capacity of the technique to accurately estimate parameters is examined for a range of scenarios.

Algorithms↗

Markov chain Monte Carlo for mapping a quantitative trait locus in outbred populations.

A Bayesian approach is presented for mapping a quantitative trait locus (QTL) using the 'Fernando and Grossman' multivariate Normal approximation to QTL inheritance. For this model, a Bayesian implementation that includes QTL position is problematic because standard Markov chain Monte Carlo (MCMC) algorithms do not mix, i.e. the QTL position gets stuck in one marker interval. This is because of the dependence of the covariance structure for the QTL effects on the adjacent markers and may be typical of the 'Fernando and Grossman' model. A relatively new MCMC technique, simulated tempering, allows mixing and so makes possible inferences about QTL position based on marginal posterior probabilities. The model was implemented for estimating variance ratios and QTL position using a continuous grid of allowed positions and was applied to simulated data of a standard granddaughter design. The results showed a smooth mixing of QTL position after implementation of the simulated tempering sampler. In this implementation, map distance between QTL and its flanking markers was artificially stretched to reduce the dependence of markers and covariance. The method generalizes easily to more complicated applications and can ultimately contribute to QTL mapping in complex, heterogeneous, human, animal or plant populations.

Animals↗

A Bayesian approach to disease gene location using allelic association.

A Bayesian approach to analysing data from family-based association studies is developed. This permits direct assessment of the range of possible values of model parameters, such as the recombination frequency and allelic associations, in the light of the data. In addition, sophisticated comparisons of different models may be handled easily, even when such models are not nested. The methodology is developed in such a way as to allow separate inferences to be made about linkage and association by including theta, the recombination fraction between the marker and disease susceptibility locus under study, explicitly in the model. The method is illustrated by application to a previously published data set. The data analysis raises some interesting issues, notably with regard to the weight of evidence necessary to convince us of linkage between a candidate locus and disease.

Alleles↗

Statistical methods for gene map construction by fluorescence in situ hybridization.

Fluorescence in situ hybridization (FISH) provides an efficient and powerful technique for ordering loci both on metaphase chromosomes and in less condensed interphase chromatin. Two-color metaphase FISH can be used to order pairs of loci relative to the centromere; two- and three-color interphase FISH can be used to accurately order trios of loci spaced within 1 Mb relative to one another. Loci separated by a distance > 1-2 Mb exhibit chromatin loops that often give rise to a statistically significant but incorrect order. We derive Bayesian methods for selecting the best locus order based on microscopic evaluation for each of these types of FISH mapping data. We then describe how the results from several two- and three-locus analyses can be combined to evaluate the approximate posterior probability of a given multilocus order within the limits of the technology utilized. These methods directly address the question of interest: What is the probability that the inferred two-, three-, or multilocus order actually is correct? We illustrate our analysis methods by applying them to previously described FISH mapping data of 14 markers in the BRCA1 region on chromosome 17q12-q21. We also propose design strategies to order a group of closely spaced (< 1 Mb) loci, two and three loci at a time, using a bisection strategy for two-color FISH data and a trisection strategy for three-color FISH data. These strategies have the best worst-case performance for ordering a new locus relative to a group of ordered loci and are nearly optimal for ordering a group of loci of unknown order. These, in conjunction with physical mapping strategies, provide efficient and reliable methods for gene map construction by FISH.

BRCA1 Protein↗

Strength of evidence for density dependence in abundance time series of 1198 species.

Population limitation is a fundamental tenet of ecology, but the relative roles of exogenous and endogenous mechanisms remain unquantified for most species. Here we used multi-model inference (MMI), a form of model averaging, based on information theory (Akaike's Information Criterion) to evaluate the relative strength of evidence for density-dependent and density-independent population dynamical models in long-term abundance time series of 1198 species. We also compared the MMI results to more classic methods for detecting density dependence: Neyman-Pearson hypothesis-testing and best-model selection using the Bayesian Information Criterion or cross-validation. Using MMI on our large database, we show that density dependence is a pervasive feature of population dynamics (median MMI support for density dependence = 74.7-92.2%), and that this holds across widely different taxa. The weight of evidence for density dependence varied among species but increased consistently, with the number of generations monitored. Best-model selection methods yielded similar results to MMI (a density-dependent model was favored in 66.2-93.9% of species time series), while the hypothesis-testing methods detected density dependence less frequently (32.6-49.8%). There were no obvious differences in the prevalence of density dependence across major taxonomic groups under any of the statistical methods used. These results underscore the value of using multiple modes of analysis to quantify the relative empirical support for a set of working hypotheses that encompass a range of realistic population dynamical behaviors.

Ecosystem↗

Algorithms for Bayesian belief-network precomputation.

Bayesian belief networks provide an intuitive and concise means of representing probabilistic relationships among the variables in expert systems. A major drawback to this methodology is its computational complexity. We present an introduction to belief networks, and describe methods for precomputing, or caching, part of a belief network based on metrics of probability and expected utility. These algorithms are examples of a general method for decreasing expected running time for probabilistic inference. We first present the necessary background, and then present algorithms for producing caches based on metrics of expected probability and expected utility. We show how these algorithms can be applied to a moderately complex belief network, and present directions for future research.

Algorithms↗

Evolutionary relationships of the cup-fungus genus Peziza and Pezizaceae inferred from multiple nuclear genes: RPB2, beta-tubulin, and LSU rDNA.

To provide a robust phylogeny of Pezizaceae, partial sequences from two nuclear protein-coding genes, RPB2 (encoding the second largest subunit of RNA polymerase II) and beta-tubulin, were obtained from 69 and 72 specimens, respectively, to analyze with nuclear ribosomal large subunit RNA gene sequences (LSU). The three-gene data set includes 32 species of Peziza, and 27 species from nine additional epigeous and six hypogeous (truffle) pezizaceous genera. Analyses of the combined LSU, RPB2, and beta-tubulin data set using parsimony, maximum likelihood, and Bayesian approaches identify 14 fine-scale lineages within Pezizaceae. Species of Peziza occur in eight of the lineages, spread among other genera of the family, confirming the non-monophyly of the genus. Although parsimony analyses of the three-gene data set produced a nearly completely resolved strict consensus tree, with increased confidence, relationships between the lineages are still resolved with mostly weak bootstrap support. Bayesian analyses of the three-gene data, however, show support for several more inclusive clades, mostly congruent with Bayesian analyses of RPB2. No strongly supported incongruence was found among phylogenies derived from the separate LSU, RPB2, and beta-tubulin data sets. The RPB2 region appeared to be the most informative single gene region based on resolution and clade support, and accounts for the greatest number of potentially parsimony informative characters within the combined data set, followed by the LSU and the beta-tubulin region. The results indicate that third codon positions in beta-tubulin are saturated, especially for sites that provide information about the deeper relationships. Nevertheless, almost all phylogenetic signal in beta-tubulin is due to third positions changes, with almost no signal in first and second codons, and contribute phylogenetic information at the "fine-scale" level within the Pezizaceae. The Pezizaceae is supported as monophyletic in analyses of the three-gene data set, but its sister-group relationships is not resolved with support. The results advocate the use of RPB2 as a marker for ascomycete phylogenetics at the inter-generic level, whereas the beta-tubulin gene appears less useful.

Ascomycota↗

An empirical comparison of information-theoretic selection criteria for multivariate behavior genetic models.

Information theory provides an attractive basis for statistical inference and model selection. However, little is known about the relative performance of different information-theoretic criteria in covariance structure modeling, especially in behavioral genetic contexts. To explore these issues, information-theoretic fit criteria were compared with regard to their ability to discriminate between multivariate behavioral genetic models under various model, distribution, and sample size conditions. Results indicate that performance depends on sample size, model complexity, and distributional specification. The Bayesian Information Criterion (BIC) is more robust to distributional misspecification than Akaike's Information Criterion (AIC) under certain conditions, and outperforms AIC in larger samples and when comparing more complex models. An approximation to the Minimum Description Length (MDL; Rissanen, J. (1996). IEEE Transactions on Information Theory 42:40-47, Rissanen, J. (2001). IEEE Transactions on Information Theory 47:1712-1717) criterion, involving the empirical Fisher information matrix, exhibits variable patterns of performance due to the complexity of estimating Fisher information matrices. Results indicate that a relatively new information-theoretic criterion, Draper's Information Criterion (DIC; Draper, 1995), which shares features of the Bayesian and MDL criteria, performs similarly to or better than BIC. Results emphasize the importance of further research into theory and computation of information-theoretic criteria.

Bayes Theorem↗

Inferring the location and effect of tumor suppressor genes by instability-selection modeling of allelic-loss data.

Cancerous tumor growth creates cells with abnormal DNA. Allelic-loss experiments identify genomic deletions in cancer cells, but sources of variation and intrinsic dependencies complicate inference about the location and effect of suppressor genes; such genes are the target of these experiments and are thought to be involved in tumor development. We investigate properties of an instability-selection model of allelic-loss data, including likelihood-based parameter estimation and hypothesis testing. By considering a special complete-data case, we derive an approximate calibration method for hypothesis tests of sporadic deletion. Parametric bootstrap and Bayesian computations are also developed. Data from three allelic-loss studies are reanalyzed to illustrate the methods.

Adenocarcinoma↗

A framework of integrating gene relations from heterogeneous data sources: an experiment on Arabidopsis thaliana.

One of the most important goals of biological investigation is to uncover gene functional relations. In this study we propose a framework for extraction and integration of gene functional relations from diverse biological data sources, including gene expression data, biological literature and genomic sequence information. We introduce a two-layered Bayesian network approach to integrate relations from multiple sources into a genome-wide functional network. An experimental study was conducted on a test-bed of Arabidopsis thaliana. Evaluation of the integrated network demonstrated that relation integration could improve the reliability of relations by combining evidence from different data sources. Domain expert judgments on the gene functional clusters in the network confirmed the validity of our approach for relation integration and network inference.

Arabidopsis↗

The position of the Hymenoptera within the Holometabola as inferred from the mitochondrial genome of Perga condei (Hymenoptera: Symphyta: Pergidae).

We sequenced most of the mitochondrial genome of the sawfly Perga condei (Insecta: Hymenoptera: Symphyta: Pergidae) and tested different models of phylogenetic reconstruction in order to resolve the position of the Hymenoptera within the Holometabola, using mitochondrial genomes. The mitochondrial genome sequenced for P. condei had less compositional bias and slower rates of molecular evolution than the honeybee, as well as a less rearranged genome organization. Phylogenetic analyses showed that, when using mitochondrial genomes, both adequate taxon sampling and more realistic models of analysis are necessary to resolve relationships among insect orders. Both parsimony and Bayesian analyses performed better when nucleotide instead of amino acid sequences were used. In particular, this study supports the placement of the Hymenoptera as sister group to the Mecopterida.

Amino Acid Sequence↗

Bayesian dynamic modeling of latent trait distributions.

Studies of latent traits often collect data for multiple items measuring different aspects of the trait. For such data, it is common to consider models in which the different items are manifestations of a normal latent variable, which depends on covariates through a linear regression model. This article proposes a flexible Bayesian alternative in which the unknown latent variable density can change dynamically in location and shape across levels of a predictor. Scale mixtures of underlying normals are used in order to model flexibly the measurement errors and allow mixed categorical and continuous scales. A dynamic mixture of Dirichlet processes is used to characterize the latent response distributions. Posterior computation proceeds via a Markov chain Monte Carlo algorithm, with predictive densities used as a basis for inferences and evaluation of model fit. The methods are illustrated using data from a study of DNA damage in response to oxidative stress.

Algorithms↗

Blind deconvolution using a variational approach to parameter, image, and blur estimation.

Following the hierarchical Bayesian framework for blind deconvolution problems, in this paper, we propose the use of simultaneous autoregressions as prior distributions for both the image and blur, and gamma distributions for the unknown parameters (hyperparameters) of the priors and the image formation noise. We show how the gamma distributions on the unknown hyperparameters can be used to prevent the proposed blind deconvolution method from converging to undesirable image and blur estimates and also how these distributions can be inferred in realistic situations. We apply variational methods to approximate the posterior probability of the unknown image, blur, and hyperparameters and propose two different approximations of the posterior distribution. One of these approximations coincides with a classical blind deconvolution method. The proposed algorithms are tested experimentally and compared with existing blind deconvolution methods.

Algorithms↗

Performance assessment for radiologists interpreting screening mammography.

When interpreting screening mammograms radiologists decide whether suspicious abnormalities exist that warrant the recall of the patient for further testing. Previous work has found significant differences in interpretation among radiologists; their false-positive and false-negative rates have been shown to vary widely. Performance assessments of individual radiologists have been mandated by the U.S. government, but concern exists about the adequacy of current assessment techniques. We use hierarchical modelling techniques to infer about interpretive performance of individual radiologists in screening mammography. While doing this we account for differences due to patient mix and radiologist attributes (for instance, years of experience or interpretive volume). We model at the mammogram level, and then use these models to assess radiologist performance. Our approach is demonstrated with data from mammography registries and radiologist surveys. For each mammogram, the registries record whether or not the woman was found to have breast cancer within one year of the mammogram; this criterion is used to determine whether the recall decision was correct. We model the false-positive rate and the false-negative rate separately using logistic regression on patient risk factors and radiologist random effects. The radiologist random effects are, in turn, regressed on radiologist attributes such as the number of years in practice. Using these Bayesian hierarchical models we examine several radiologist performance metrics. The first is the difference between the false-positive or false-negative rate of a particular radiologist and that of a hypothetical 'standard' radiologist with the same attributes and the same patient mix. A second metric predicts the performance of each radiologist on hypothetical mammography exams with particular combinations of patient risk factors (which we characterize as 'typical', 'high-risk', or 'low-risk'). The second metric can be used to compare one radiologist to another, while the first metric addresses how the radiologist is performing compared to an appropriate standard. Interval estimates are given for the metrics, thereby addressing uncertainty. The particular novelty in our contribution is to estimate multiple performance rates (sensitivity and specificity). One can even estimate a continuum of performance rates such as a performance curve or ROC curve using our models and we describe how this may be done. In addition to assessing radiologists in the original data set, we also show how to infer about the performance of a new radiologist with new case mix, new outcome data, and new attributes without having to refit the model.

Adult↗

Comparing alignment methods for inferring the history of the new world lizard genus Mabuya (Squamata: Scincidae).

The rapid increase in the ability to generate molecular data, and the focus on model-based methods for tree reconstruction have greatly advanced the use of phylogenetics in many fields. The recent flurry of new analytical techniques has focused almost solely on tree reconstruction, whereas alignment issues have received far less attention. In this paper, we use a diverse sampling of gene regions from lizards of the genus Mabuya to compare the impact, on phylogeny estimation, of new maximum likelihood alignment algorithms with more widely used methods. Sequences aligned under different optimality criteria are analyzed using partitioned Bayesian analysis with independent models and parameter settings for each gene region, and the most strongly supported phylogenetic hypothesis is then used to test the hypothesis of two colonizations of the New World by African scincid lizards. Our results show that the consistent use of model-based methods in both alignment and tree reconstruction leads to trees with more optimal likelihood scores than the use of independent criteria in alignment and tree reconstruction. We corroborate and extend earlier evidence for two independent colonizations of South America by scincid lizards. Relationships within South American Mabuya are found to be in need of taxonomic revision, specifically complexes under the names M. heathi, M. agilis, and M. bistriata (sensu, M.T. Rodrigues, Papeis Avulsos de Zoologia 41 (2000) 313).

Animals↗

Phylogeny and biogeography of the Petaurista philippensis complex (Rodentia: Sciuridae), inter- and intraspecific relationships inferred from molecular and morphometric analysis.

With modified DNA extraction and purification protocols, the complete cytochrome b gene sequences (1140 bp) were determined from degraded museum specimens. Molecular analysis and morphological examination of cranial characteristics of the giant flying squirrels of Petaurista philippensis complex (P. grandis, P. hainana, and P. yunanensis) and other Petaurista species yielded new insights into long-standing controversies in the Petaurista systematics. Patterns of genetic variations and morphological differences observed in this study indicate that P. hainana, P. albiventer, and P. yunanensis can be recognized as distinct species, and P. grandis and P. petaurista are conspecific populations. Phylogenetic relationships reconstructed by using parsimony, likelihood, and Bayesian methods reveal that, with P. leucogenys as the basal branch, all Petaurista groups formed two distinct clades. Petaurista philippensis, P. hainana, P. yunanensis, and P. albiventer are clustered in the same clade, while P. grandis shows a close relationship to P. petaurista. Deduced divergence times based on Bayesian analysis and the transversional substitution at the third codon suggest that the retreating of glaciers and upheavals or movements of tectonic plates in the Pliocene-Pleistocene were the major factors responsible for the present geographical distributions of Petaurista groups.

Animals↗