Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,342 records · Page 75Linked to original sources

A Bayesian hierarchical model for accident and injury surveillance.

This article presents a recent study which applies Bayesian hierarchical methodology to model and analyse accident and injury surveillance data. A hierarchical Poisson random effects spatio-temporal model is introduced and an analysis of inter-regional variations and regional trends in hospitalisations due to motor vehicle accident injuries to boys aged 0-24 in the province of British Columbia, Canada, is presented. The objective of this article is to illustrate how the modelling technique can be implemented as part of an accident and injury surveillance and prevention system where transportation and/or health authorities may routinely examine accidents, injuries, and hospitalisations to target high-risk regions for prevention programs, to evaluate prevention strategies, and to assist in health planning and resource allocation. The innovation of the methodology is its ability to uncover and highlight important underlying structure of the data. Between 1987 and 1996, British Columbia hospital separation registry registered 10,599 motor vehicle traffic injury related hospitalisations among boys aged 0-24 who resided in British Columbia, of which majority (89%) of the injuries occurred to boys aged 15-24. The injuries were aggregated by three age groups (0-4, 5-14, and 15-24), 20 health regions (based of place-of-residence), and 10 calendar years (1987 to 1996) and the corresponding mid-year population estimates were used as 'at risk' population. An empirical Bayes inference technique using penalised quasi-likelihood estimation was implemented to model both rates and counts, with spline smoothing accommodating non-linear temporal effects. The results show that (a) crude rates and ratios at health region level are unstable, (b) the models with spline smoothing enable us to explore possible shapes of injury trends at both the provincial level and the regional level, and (c) the fitted models provide a wealth of information about the patterns (both over space and time) of the injury counts, rates and ratios. During the 10-year period, high injury risk ratios evolved from northwest to central-interior and the southeast [corrected].

Accidents, Traffic↗

On flexible finite polygenic models for multiple-trait evaluation.

Finite polygenic models (FPM) might be an alternative to the infinitesimal model (TIM) for the genetic evaluation of pedigreed multiple-generation populations for multiple quantitative traits. I present a general flexible Bayesian method that includes the number of genes in the FPM as an additional random variable. Markov-chain Monte Carlo techniques such as Gibbs sampling and the reversible jump sampler are used for implementation. Sampling of genotypes of all genes in the FPM is done via the use of segregation indicators. A broad range of FPM models, some combined with TIM, are empirically tested for the estimation of variance components and the number of genes in the FPM. Four simulation scenarios were studied, including genetic models with 5 or 50 additive independent diallelic genes affecting the traits, and random selection or selection on one of the traits was performed. The results in this study were based on ten replicates per simulation scenario. In the case of random selection, uniform priors on additive gene effects led to posterior mean estimates of genetic variance that were positively correlated with the number of genes fitted in the FPM. In the case of trait selection, assuming normal priors on gene effects also led to genetic variance estimates for the selected trait that were negatively correlated with the number of genes in the FPM. This negative correlation was not observed for the unselected trait. Treating the number of genes in the FPM as random revealed a positive correlation between prior and posterior mean estimates of this number, but the prior hardly affected the posterior estimates of genetic variance. Posterior inferences about the number of genes should be considered to be indicative where trait selection seems to improve the power of distinguishing between TIM and FPM. Based on the results of this study, I suggest not replacing TIM by the FPM, but combining TIM and FPM with the number of genes treated as random, to facilitate a highly flexible and thereby robust method for variance component estimation in pedigreed populations. Further study is required to explore the full potential of these models under different genetic model assumptions.

Analysis of Variance↗

Evidence for multiple reversals of asymmetric mutational constraints during the evolution of the mitochondrial genome of metazoa, and consequences for phylogenetic inferences.

Mitochondrial DNA (mtDNA) sequences are comonly used for inferring phylogenetic relationships. However, the strand-specific bias in the nucleotide composition of the mtDNA, which is thought to reflect assymetric mutational constraints, combined with the important compositional heterogeneity among taxa, are known to be highly problematic for phylogenetic analyses. Here, nucleotide composition was compared across 49 species of Metazoa (34 arthropods, 2 annelids, 2 molluscs, and 11 deuterosomes), and analyzed for a mtDNA fragment including six protein-coding genes, i.e., atp6, atp8, cox1, cox2, cox3, and nad2. The analyses show that most metazoan species present a clear strand assymetry, where one strand is biased in favor of A and C, whereas the other strand has reverse bias, i.e. in favor of T and G. the origin of this strand bias can be related to assymetric mutational constraints involving deaminations of A and C nucleotides during the replication and/or transcription processes. The analyses reveal that six unrelated genera are characterized by a reversal of the usual strand bias, i.e., Argiope (Araneae), Euscorpius (Scorpiones), Tigrioupus (Maxillopoda), Branchiostoma (Cephalochordata) Florometra (Echinodermata), and Katharina (Mollusca). It is proposed that assymetric mutational constraints have been independantly reversed in these six genera, through an inversion of the control region, i.e., the region that contains most regulatory elements for replication and transcription of the mtDNA. We show that reversals of assymetric mutational constraints have dramatic consequences on the phylogenetic analyses, as taxa characterized by reverse strand bias tend to group together due to long-branch attraction artifacts. We propose a new method for limiting this specific problem in tree reconstruction under the Bayesian approach. We apply our method to deal with the question of phylogenetic relationships of the major lineages of Arthropoda, This new approach provides a better congruence with nuclear analyses based on mtDNA sequences, our data suggest that Chelicerata, Crustacea, Myriapoda, Pancrustacea, and Paradoxopoda are monophyletic.

Animals↗

How vague is vague? A simulation study of the impact of the use of vague prior distributions in MCMC using WinBUGS.

There has been a recent growth in the use of Bayesian methods in medical research. The main reasons for this are the development of computer intensive simulation based methods such as Markov chain Monte Carlo (MCMC), increases in computing power and the introduction of powerful software such as WinBUGS. This has enabled increasingly complex models to be fitted. The ability to fit these complex models has led to MCMC methods being used as a convenient tool by frequentists, who may have no desire to be fully Bayesian. Often researchers want 'the data to dominate' when there is no prior information and thus attempt to use vague prior distributions. However, with small amounts of data the use of vague priors can be problematic. The results are potentially sensitive to the choice of prior distribution. In general there are fewer problems with location parameters. The main problem is with scale parameters. With scale parameters, not only does one have to decide the distributional form of the prior distribution, but also whether to put the prior distribution on the variance, standard deviation or precision. We have conducted a simulation study comparing the effects of 13 different prior distributions for the scale parameter on simulated random effects meta-analysis data. We varied the number of studies (5, 10 and 30) and compared three different between-study variances to give nine different simulation scenarios. One thousand data sets were generated for each scenario and each data set was analysed using the 13 different prior distributions. The frequentist properties of bias and coverage were investigated for the between-study variance and the effect size. The choice of prior distribution was crucial when there were just five studies. There was a large variation in the estimates of the between-study variance for the 13 different prior distributions. With a large number of studies the choice of prior distribution was less important. The effect size estimated was not biased, but the precision with which it was estimated varied with the choice of prior distribution leading to varying coverage intervals and, potentially, to different statistical inferences. Again there was less of a problem with a larger number of studies. There is a particular problem if the between-study variance is close to the boundary at zero, as MCMC results tend to produce upwardly biased estimates of the between-study variance, particularly if inferences are based on the posterior mean. The choice of 'vague' prior distribution can lead to a marked variation in results, particularly in small studies. Sensitivity to the choice of prior distribution should always be assessed.

Anti-Bacterial Agents↗

A flexible framework for robust and efficient Mendelian randomization with debiasing.

Mendelian randomization (MR) has been widely used to infer causal relationships between exposures and outcomes in epidemiological studies. However, classical MR assumptions can be violated when genetic variants are associated with outcomes through pathways other than the exposure, leading to uncorrelated and/or correlated pleiotropy. Additionally, measurement error arising from the inherent uncertainty in summary statistics obtained from large-scale genome-wide association studies can introduce bias into the causal effect estimate. To address these issues, we develop a debiased mixture inverse variance weighting ($\mathsf{dmIVW}$) method with three major advantages. First, it is capable of simultaneously handling various types of pleiotropy and eliminating the bias caused by uncertainty. Second, it can guard against distortion caused by invalid genetic variants while effectively harnessing their information. Third, our unified framework facilitates a fair comparison and combination of a series of submodels, encompassing several popular MR methods as special cases. Through real data applications, the effectiveness and robustness of $\mathsf{dmIVW}$ in estimating the causal effects of risk factors on common diseases are demonstrated.

Mendelian Randomization Analysis↗

A mitogenomic timescale for birds detects variable phylogenetic rates of molecular evolution and refutes the standard molecular clock.

Current understanding of the diversification of birds is hindered by their incomplete fossil record and uncertainty in phylogenetic relationships and phylogenetic rates of molecular evolution. Here we performed the first comprehensive analysis of mitogenomic data of 48 vertebrates, including 35 birds, to derive a Bayesian timescale for avian evolution and to estimate rates of DNA evolution. Our approach used multiple fossil time constraints scattered throughout the phylogenetic tree and accounts for uncertainties in time constraints, branch lengths, and heterogeneity of rates of DNA evolution. We estimated that the major vertebrate lineages originated in the Permian; the 95% credible intervals of our estimated ages of the origin of archosaurs (258 MYA), the amniote-amphibian split (356 MYA), and the archosaur-lizard divergence (278 MYA) bracket estimates from the fossil record. The origin of modern orders of birds was estimated to have occurred throughout the Cretaceous beginning about 139 MYA, arguing against a cataclysmic extinction of lineages at the Cretaceous/Tertiary boundary. We identified fossils that are useful as time constraints within vertebrates. Our timescale reveals that rates of molecular evolution vary across genes and among taxa through time, thereby refuting the widely used mitogenomic or cytochrome b molecular clock in birds. Moreover, the 5-Myr divergence time assumed between 2 genera of geese (Branta and Anser) to originally calibrate the standard mitochondrial clock rate of 0.01 substitutions per site per lineage per Myr (s/s/l/Myr) in birds was shown to be underestimated by about 9.5 Myr. Phylogenetic rates in birds vary between 0.0009 and 0.012 s/s/l/Myr, indicating that many phylogenetic splits among avian taxa also have been underestimated and need to be revised. We found no support for the hypothesis that the molecular clock in birds "ticks" according to a constant rate of substitution per unit of mass-specific metabolic energy rather than per unit of time, as recently suggested. Our analysis advances knowledge of rates of DNA evolution across birds and other vertebrates and will, therefore, aid comparative biology studies that seek to infer the origin and timing of major adaptive shifts in vertebrates.

Animals↗

Building phenotypic character matrices for phylogenetic inference: exploration of 35 years of practice.

Recent methodological development in phylogenetic inference has focused predominantly on molecular data. However, renewed interest in other data types, particularly morphological data, has followed from the increased recognition of the power of total evidence and tip-dating approaches, including fossil data, for inference of time-scaled trees and rates of evolution. However, attention has largely focused on the improvement of models of morphological evolution and other analytical tools with much less discussion about data acquisition itself. Here we review past and current practice for describing and collecting morphological data for phylogenetic inference. We present a systematic review of 164 phylogenetic analyses conducted over the last 35 years and focused on a diverse group of extinct arthropods: trilobites. Trends in increasing matrix size, data type, and coding strategy are evident. Where present, polymorphic characters have been predominantly derived from discretized continuous characters, although increasingly practitioners are utilizing alternative approaches for the treatment of quantitative characters. Not surprisingly, traditional indices that describe character consistency are highly correlated with matrix size but show surprising variation at different taxonomic scales. More recent attempts to describe data quality using information theory imply that characters can have high information content even if data are missing for many tips, providing support against the exclusion of characters because of missing data. In consideration of this, as well as advances in the study of developmental biology and variational complexity, we identify several avenues for increasing the quality and quantity of morphological data going forward.

Phylogeny↗

Replicating lipid micelles: a feasible precursor to the origin of life and the earliest appearance of genomes.

The most commonly accepted scenario of early Earth includes: creation of the universe around 13.8 Ga (Giga-annus; or 109 years ago); establishment of our solar system ~ 4.60 Ga; and formation of Earth ~ 4.54 Ga. The earliest life forms on our planet so far observed to have existed, are microbes that left signals of their presence in rocks ~ 3.6 Ga - suggesting that Life forms existed within the first 940 million years after Earth's formation. However, an intriguing recent publication [1] infers that the last universal common ancestor (LUCA) likely existed by 4.2 Ga, and that the inferred LUCA had a genome of at least 2.5 Mb of DNA, encoding around 2,600 proteins; this suggests that sophisticated Life might have existed within the first 340 million years after Earth was formed. The commonly accepted geological history of early Earth suggests that the turbulent Hadean Eon lasted until 4.0 Ga, with the Late Heavy Bombardment (LHB) period occurring around 4.1 to 3.8 Ga. If Earth during the Hadean exhibited a molten surface, intense volcanic activity, and constant bombardment by asteroids and comets - how were sensitive molecules (e.g., nucleic acids, proteins) able to survive? Considering the "Lipid First" hypothesis [2], we propose that replicating lipid micelles are feasible candidates for having populated much of Earth's deep hydrothermal vents and turbulent surface within the first 340 million years of Earth's existence. These lipid micelles could therefore have provided a plausible form of "protective capsules" inside which early Life's sensitive molecules were able to evolve.

Origin of Life↗

Dating dispersal and radiation in the gymnosperm Gnetum (Gnetales)--clock calibration when outgroup relationships are uncertain.

Most implementations of molecular clocks require resolved topologies. However, one of the Bayesian relaxed clock approaches accepts input topologies that include polytomies. We explored the effects of resolved and polytomous input topologies in a rate-heterogeneous sequence data set for Gnetum, a member of the seed plant lineage Gnetales. Gnetum has 10 species in South America, 1 in tropical West Africa, and 20 to 25 in tropical Asia, and explanations for the ages of these disjunctions involve long-distance dispersal and/or the breakup of Gondwana. To resolve relationships within Gnetum, we sequenced most of its species for six loci from the chloroplast (rbcL, matK, and the trnT-trnF region), the nucleus (rITS/5.8S and the LEAFY gene second intron), and the mitochondrion (nad1 gene second intron). Because Gnetum has no fossil record, we relied on fossils from other Gnetales and from the seed plant lineages conifers, Ginkgo, cycads, and angiosperms to constrain a molecular clock and obtain absolute times for within-Gnetum divergence events. Relationships among Gnetales and the other seed plant lineages are still unresolved, and we therefore used differently resolved topologies, including one that contained a basal polytomy among gymnosperms. For a small set of Gnetales exemplars (n = 13) in which rbcL and matK satisfied the clock assumption, we also obtained time estimates from a strict clock, calibrated with one outgroup fossil. The changing hierarchical relationships among seed plants (and accordingly changing placements of distant fossils) resulted in small changes of within-Gnetum estimates because topologically closest constraints overrode more distant constraints. Regardless of the seed plant topology assumed, relaxed clock estimates suggest that the extant clades of Gnetum began diverging from each other during the Upper Oligocene. Strict clock estimates imply a mid-Miocene divergence. These estimates, together with the phylogeny for Gnetum from the six combined data sets, imply that the single African species of Gnetum is not a remnant of a once Gondwanan distribution. Miocene and Pliocene range expansions are inferred for the Asian subclades of Gnetum, which stem from an ancestor that arrived from Africa. These findings fit with seed dispersal by water in several species of Gnetum, morphological similarities among apparently young species, and incomplete concerted evolution in the nuclear ITS region.

Base Sequence↗

Identifying conflicting signal in a multigene analysis reveals a highly resolved tree: the phylogeny of Rodentia (Mammalia).

Homoplasy among morphological characters has hindered inference of higher level rodent phylogeny for over 100 years. Initial molecular studies, based primarily on single genes, likewise produced little resolution of the deep relationships among rodent families. Two recent molecular studies (Huchon et al., 2002, Mol. Biol. Evol. 19:1053-1065; Adkins et al., 2003, Mol. Phylogenet. Evol. 26:409-420), using larger samples from the nuclear genome, have produced phylogenies that are generally concordant with each other, but many of the deep superfamilial nodes were still lacking substantial statistical support. Data are presented here for a total of approximately 3,600 base pairs from portions of three different nuclear protein-coding genes, CB1, IRBP, and RAG2, from 19 rodents and three outgroups. Separate analyses, with data partitioned according to both genes and codon position, produced conflicting results. Trees obtained from all partitions of CB1 and RAG2 and those obtained from the first- plus second-position sites of IRBP were generally concordant with each other and the trees from the two recent studies, whereas trees obtained from the third-position sites of IRBP were not. Although the IRBP third-position sites represent only 1/9 of the total data set, combined analyses using either parsimony or likelihood resulted in trees in agreement with the IRBP third-position sites and in disagreement with the remaining 8/9 of the sites from this data set and the two recent multigene studies. In contrast, maximum-likelihood analysis using a site-specific rates model did recover a tree that is highly congruent with the trees in the two recent studies. If the IRBP third-position sites are removed from the current data set, then combined likelihood analyses obtain a tree that is highly congruent with those of the two recent studies. This analysis also provides, for the first time in a study of rodent phylogeny, robust statistical support for every bipartition, with just one exception. This tree divides rodents into two major clades. The first contains Myodonta (Muroidea plus Dipodidae) and the only unresolved trichotomy, from which descend Geomyoidea, Pedetidae, and Castoridae. On the other side of the root is a clade containing Sciuroidea plus Gliridae, and Hystricognathi. Some uncertainty remains on the placement of the root. Trees on which the Hystricognathi are the basal sister group to Myodonta, Geomyoidea, Pedetidae, and Castoridae are also found within a Bayesian 95% credible set, as estimated by Metropolis-coupled Markov chain Monte Carlo sampling.

Animals↗