Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian dating”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Primate molecular divergence dates.

With genomic data, alignments can be assembled that greatly increase the number of informative sites for analysis of molecular divergence dates. Here, we present an estimate of the molecular divergence dates for all of the major primate groups. These date estimates are based on a Bayesian analysis of approximately 59.8 kbp of genomic data from 13 primates and 6 mammalian outgroups, using a range of paleontologically supported calibration estimates. Results support a Cretaceous last common ancestor of extant primates (approximately 77 mya), an Eocene divergence between platyrrhine and catarrhine primates (approximately 43 mya), an Oligocene origin of apes and Old World monkeys (approximately 31 mya), and an early Miocene (approximately 18 mya) divergence of Asian and African great apes. These dates are examined in the context of other molecular clock studies.

Animals↗

Catarrhine primate divergence dates estimated from complete mitochondrial genomes: concordance with fossil and nuclear DNA evidence.

Accurate divergence date estimates improve scenarios of primate evolutionary history and aid in interpretation of the natural history of disease-causing agents. While molecule-based estimates of divergence dates of taxa within the superfamily Hominoidea (apes and humans) are common in the literature, few such estimates are available for the Cercopithecoidea (Old World monkeys), the sister taxon of the hominoids in the primate infraorder Catarrhini. To help fill this gap, we have sequenced the entire mitochondrial DNA (mtDNA) genomes from a representative of three cercopithecoid tribes, Cercopithecini (Chlorocebus aethiops), Colobini (Colobus guereza), and Presbytini (Trachypithecus obscurus), and analyzed these new data together with other catarrhine mtDNA genomes available in public databases. Molecular divergence date estimates are dependent on calibration points gleaned from the paleontological record. We defined criteria for the selection of good calibration points and identified three points meeting these criteria: Homo-Pan, 6.0 Ma; Pongo-hominines, 14.0 Ma; hominoid/cercopithecoid, 23.0 Ma. Because a uniform molecular clock does not fit the catarrhine mtDNA data, we estimated divergence dates using a penalized likelihood and a Bayesian method, both of which take into account the effects of rate differences on lineages, phylogenetic tree structure, and multiple calibration points. The penalized likelihood method applied to the coding regions of the mtDNA genome yielded the following divergence date estimates, with approximate 95% confidence intervals: cercopithecine-colobine, 16.2 (14.4-17.9) Ma; colobin-presbytin, 10.9 (9.6-12.3) Ma; cercopithecin-papionin, 11.6 (10.3-12.9) Ma; and Macaca-Papio, 9.8 (8.6-10.9) Ma. Within the hominoids, the following dates were inferred: hylobatid-hominid, 16.8 (15.0-18.5) Ma; Gorilla-Homo+Pan, 8.1 (7.1-9.0) Ma; Pongo pygmaeus pygmaeus-P. p. abelii, 4.1 (3.5-4.7) Ma; and Pan troglodytes-P. paniscus, 2.4 (2.0-2.7) Ma. These dates were similar to those found using penalized likelihood on other subsets of the data, but slightly younger than several of the Bayesian estimates.

Africa↗

Bayesian Inference of Pathogen Phylogeography using the Structured Coalescent Model.

Over the past decade, pathogen genome sequencing has become well established as a powerful approach to study infectious disease epidemiology. In particular, when multiple genomes are available from several geographical locations, comparing them is informative about the relative size of the local pathogen populations as well as past migration rates and events between locations. The structured coalescent model has a long history of being used as the underlying process for such phylogeographic analysis. However, the computational cost of using this model does not scale well to the large number of genomes frequently analysed in pathogen genomic epidemiology studies. Several approximations of the structured coalescent model have been proposed, but their effects are difficult to predict. Here we show how the exact structured coalescent model can be used to analyse a precomputed dated phylogeny, in order to perform Bayesian inference on the past migration history, the effective population sizes in each location, and the directed migration rates from any location to another. We describe an efficient reversible jump Markov Chain Monte Carlo scheme which is implemented in a new R package StructCoalescent. We use simulations to demonstrate the scalability and correctness of our method and to compare it with existing software. We also applied our new method to several state-of-the-art datasets on the population structure of real pathogens to showcase the relevance of our method to current data scales and research questions.

Bayes Theorem↗

Evaluating Neanderthal genetics and phylogeny.

The retrieval of Neanderthal (Homo neanderthalsensis) mitochondrial DNA is thought to be among the most significant ancient DNA contributions to date, allowing conflicting hypotheses on modern human (Homo sapiens) evolution to be tested directly. Recently, however, both the authenticity of the Neanderthal sequences and their phylogenetic position outside contemporary human diversity have been questioned. Using Bayesian inference and the largest dataset to date, we find strong support for a monophyletic Neanderthal clade outside the diversity of contemporary humans, in agreement with the expectations of the Out-of-Africa replacement model of modern human origin. From average pairwise sequence differences, we obtain support for claims that the first published Neanderthal sequence may include errors due to postmortem damage in the template molecules for PCR. In contrast, we find that recent results implying that the Neanderthal sequences are products of PCR artifacts are not well supported, suffering from inadequate experimental design and a presumably high percentage (>68%) of chimeric sequences due to "jumping PCR" events.

Animals↗

Molecular phylogeny and biogeography of an ancient Holarctic lineage of mygalomorph spiders (Araneae: Antrodiaetidae: Antrodiaetus).

The mygalomorph spider genera Antrodiaetus and Atypoides (Antrodiaetidae) belong to an ancient lineage that has persisted since at least the Cretaceous. These spiders display a classic disjunct Holarctic distribution with species in the eastern Palaearctic plus the western and eastern Nearctic. Prior phylogenetic analyses of this group have been proposed on the basis of morphology, but lack strong support and independent corroboration. Here we present the first phylogenetic analysis of species-level relationships based on molecular data obtained from the mitochondrial (cytochrome c oxidase subunit I) and nuclear (18S and 28S rRNA) genomes. Analyses corroborate earlier findings that Atypoides forms a paraphyletic grade with respect to Antrodiaetus, and consequently, that genus is formally synonymized under Antrodiaetus. In addition, our results support the relatively early divergence of Antrodiaetus roretzi. Antrodiaetus pacificus is "paraphyletic" with respect to the A. lincolnianus group and is likely an assemblage of numerous species. The final topology based on a combined molecular dataset, in conjunction with two different molecular dating techniques (penalized likelihood plus a Bayesian approach) and ancestral distribution reconstructions, was used to infer the historical biogeography of these spiders. Trans-Beringian and trans-Atlantic routes appear to account for the present-day distribution of Antrodiaetus in Japan and North America. Future studies on Antrodiaetus phylogeny will be used to address questions regarding morphological stasis and the evolution of quantitative morphological characters.

Animals↗

Chronology for the Aegean Late Bronze Age 1700-1400 B.C.

Radiocarbon (carbon-14) data from the Aegean Bronze Age 1700-1400 B.C. show that the Santorini (Thera) eruption must have occurred in the late 17th century B.C. By using carbon-14 dates from the surrounding region, cultural phases, and Bayesian statistical analysis, we established a chronology for the initial Aegean Late Bronze Age cultural phases (Late Minoan IA, IB, and II). This chronology contrasts with conventional archaeological dates and cultural synthesis: stretching out the Late Minoan IA, IB, and II phases by approximately 100 years and requiring reassessment of standard interpretations of associations between the Egyptian and Near Eastern historical dates and phases and those in the Aegean and Cyprus in the mid-second millennium B.C.

Archaeology↗

Predicting protein secondary structure using neural net and statistical methods.

A comparison of neural network methods and Bayesian statistical methods is presented for prediction of the secondary structure of proteins given their primary sequence. The Bayesian method makes the unphysical assumption that the probability of an amino acid occurring in each position in the protein is independent of the amino acids occurring elsewhere. However, we find the predictive accuracy of the Bayesian method to be only minimally less than the accuracy of the most sophisticated methods used to date. We present the relationship of neural network methods to Bayesian statistical methods and show that, in principle, neural methods offer considerable power, although apparently they are not particularly useful for this problem. In the process, we derive a neural formalism in which the output neurons directly represent the conditional probabilities of structure class. The probabilistic formalism allows introduction of a new objective function, the mutual information, which translates the notion of correlation as a measure of predictive accuracy into a useful training measure. Although a similar accuracy to other approaches (utilizing a mean-square error) is achieved using this new measure, the accuracy on the training set is significantly and tantalizingly higher, even though the number of adjustable parameters remains the same. The mutual information measure predicts a greater fraction of helix and sheet structures correctly than the mean-square error measure, at the expense of coil accuracy, precisely as it was designed to do. By combining the two objective functions, we obtain a marginally improved accuracy of 64.4%, with Matthews coefficients C alpha, C beta and Ccoil of 0.40, 0.32 and 0.42, respectively. However, since all methods to date perform only slightly better than the Bayes algorithm, which entails the drastic assumption of independence of amino acids, one is forced to conclude that little progress has been made on this problem, despite the application of a variety of sophisticated algorithms such as neural networks, and that further advances will require a better understanding of the relevant biophysics.

Bayes Theorem↗

Unraveling the evolutionary radiation of the thoracican barnacles using molecular and morphological evidence: a comparison of several divergence time estimation approaches.

The Thoracica includes the ordinary barnacles found along the sea shore and is the most diverse and well-studied superorder of Cirripedia. However, although the literature abounds with scenarios explaining the evolution of these barnacles, very few studies have attempted to test these hypotheses in a phylogenetic context. The few attempts at phylogenetic analyses have suffered from a lack of phylogenetic signal and small numbers of taxa. We collected DNA sequences from the nuclear 18S, 28S, and histone H3 genes and the mitochondrial 12S and 16S genes (4,871 bp total) and data for 37 adult and 53 larval morphological characters from 43 taxa representing all the extant thoracican suborders (except the monospecific Brachylepadomorpha). Four Rhizocephala (highly modified parasitic barnacles) taxa and a Rhizocephala + Acrothoracica (burrowing barnacles) hypothetical ancestor were used as the outgroup for the molecular and morphological analyses, respectively. We analyzed these data separately and combined using maximum likelihood (ML) under "hill-climbing" and genetic algorithm heuristic searches, maximum parsimony procedures, and Bayesian inference coupled with Markov chain Monte Carlo techniques under mixed and homogeneous models of nucleotide substitution. The resulting phylogenetic trees answered key questions in barnacle evolution. The four-plated Iblomorpha were shown as the most primitive thoracican, and the plateless Heteralepadomorpha were placed as the sister group of the Lepadomorpha. These relationships suggest for the first time in an invertebrate that exoskeleton biomineralization may have evolved from phosphatic to calcitic. Sessilia (nonpedunculate) barnacles were depicted as monophyletic and appear to have evolved from a stalked (pedunculate) multiplated (5+) scalpelloidlike ancestor rather than a five-plated lepadomorphan ancestor. The Balanomorpha (symmetric sessile barnacles) appear to have the following relationship: (Chthamaloidea(Coronuloidea(Tetraclitoidea, Balanoidea))). Thoracican divergence times were estimated under ML-based local clock, Bayesian, and penalized likelihood approaches using an 18S data set and three calibration points: Heteralepadomorpha = 530 million years ago (MYA), Scalpellomorpha = 340 MYA, and Verrucomorpha = 120 MYA. Estimated dates varied considerably within and between approaches depending on the calibration point. Highly parameterized local clock models that assume independent rates (r > or = 15) for confamilial or congeneric species generated the most congruent estimates among calibrations and agreed more closely with the barnacle fossil record. Reasonable estimates were also obtained under the Bayesian procedure of Kishino et al. (2001, Mol. Biol. Evol. 18:352-361) but using multiple calibrations. Most of the dates estimated under the Bayesian procedure of Aris-Brosou and Yang (2002, Syst. Biol. 51:703-714) and the penalized likelihood method using single and/or multiple calibrations were inconsistent among calibrations and did not fit the fossil record.

Animals↗

Predicting event times in clinical trials when treatment arm is masked.

Because power is primarily determined by the number of events in event-based clinical trials, the timing for interim or final analysis of data is often determined based on the accrual of events during the course of the study. Thus, it is of interest to predict early and accurately the time of a landmark interim or terminating event. Existing Bayesian methods may be used to predict the date of the landmark event, based on current enrollment, event, and loss to follow-up, if treatment arms are known. This work extends these methods to the case where the treatment arms are masked by using a parametric mixture model with a known mixture proportion. Posterior simulation using the mixture model is compared with methods assuming a single population. Comparison of the mixture model with the single-population approach shows that with few events, these approaches produce substantially different results and that these results converge as the prediction time is closer to the landmark event. Simulations show that the mixture model with diffuse priors can have better coverage probabilities for the prediction interval than the nonmixture models if a treatment effect is present.

Bayes Theorem↗

Origin of mitochondrial DNA diversity of domestic yaks.

BACKGROUND: The domestication of plants and animals was extremely important anthropologically. Previous studies have revealed a general tendency for populations of livestock species to include deeply divergent maternal lineages, indicating that they were domesticated in multiple, independent events from genetically discrete wild populations. However, in water buffalo, there are suggestions that a similar deep maternal bifurcation may have originated from a single population. These hypotheses have rarely been rigorously tested because of a lack of sufficient wild samples. To investigate the origin of the domestic yak (Poephagus grunnies), we analyzed 637 bp of maternal inherited mtDNA from 13 wild yaks (including eight wild yaks from a small population in west Qinghai) and 250 domesticated yaks from major herding regions. RESULTS: The domestic yak populations had two deeply divergent phylogenetic groups with a divergence time of > 100,000 yrs BP. We here show that haplotypes clustering with two deeply divergent maternal lineages in domesticated yaks occur in a single, small, wild population. This finding suggests that all domestic yaks are derived from a single wild gene pool. However, there is no clear correlation of the mtDNA phylogenetic clades and the 10 morphological types of sampled yaks indicating that the latter diversified recently. Relatively high diversity was found in Qinghai and Tibet around the current wild distribution, in accordance with previous suggestions that the earliest domestications occurred in this region. Conventional molecular clock estimation led to an unrealistic early dating of the start of the domestication. However, Bayesian estimation of the coalescence time allowing a relaxation of the mutation rate are better in agreement with a domestication during the Holocene as supported by archeological records. CONCLUSION: The information gathered here and the previous studies of other animals show that the demographic histories of domestication of livestock species were highly diverse despite the common general feature of deeply divergent maternal lineages. The results further suggest that domestication of local wild prey ungulate animals was a common occurrence during the development of human civilization following the postglacial colonization in different locations of the world, including the high, arid Qinghai-Tibetan Plateau.

Animals↗

Phylogenetic reconstruction of a known HIV-1 CRF04_cpx transmission network using maximum likelihood and Bayesian methods.

The CRF04_cpx strains of HIV-1 accounts for approximately 2-10% of the infected population in Greece, across different transmission risk groups. CRF04_cpx was the lineage documented in an HIV-1 transmission network in Thessalonica, northern Greece. Most of the transmissions occurred through unprotected heterosexual contacts between 1989 and 1993. Blood samples were available for six patients, obtained 6-10 years later, except for one patient sampled in 1991. Our objective was to examine whether the transmission history is compatible with the evolutionary tree of the virus, in partial gag, partial env, and partial gag+env. The inferred phylogenetic tree obtained using maximum likelihood and Bayesian methods in partial gag+env was much closer to the transmission tree than that using either env or gag separately. Our findings suggest that the epidemiological relationships among patients who have been infected by a common source correspond almost exactly to the evolutionary trees of the virus, given that enough phylogenetic signal is present in the alignment. Moreover, we found evidence that recombination is not the most parsimonious explanation for the phylogenetic incongruence between gag and env. For patients with known infection dates, the estimated dates of the coalescent events obtained using molecular clock calculations based on a newly developed Bayesian method in gag + env were in agreement with the actual infection dates.

Bayes Theorem↗

Estimating the rate of evolution of the rate of molecular evolution.

A simple model for the evolution of the rate of molecular evolution is presented. With a Bayesian approach, this model can serve as the basis for estimating dates of important evolutionary events even in the absence of the assumption of constant rates among evolutionary lineages. The method can be used in conjunction with any of the widely used models for nucleotide substitution or amino acid replacement. It is illustrated by analyzing a data set of rbcL protein sequences.

Algorithms↗

Taxon sampling effects in molecular clock dating: an example from the African Restionaceae.

Three commonly used molecular dating methods for correction of variable rates (non-parametric rate smoothing, penalized likelihood, and Bayesian rate correction) as well as the assumption of a global molecular clock were tested for sensitivity to taxon sampling. The test dataset of 6854 basepairs for 300 terminals includes a nearly complete sample of the Restio-clade of the African Restionaceae (272 of the 288 species), as well as 26 outgroup species. Of this, nested subsets of 35, 51, 80, 120, 150, and the full 300 species were used. Molecular dating experiments with these datasets showed that all methods are sensitive to undersampling, but that this effect is more severe in analyses that use more extreme rate smoothing. Additionally, the undersampling effect is positively related to distance from the calibration node. The combined effect of undersampling and distance from the calibration node resulted in up to threefold differences in the age estimation of nodes from the same dataset with the same calibration point. We suggest that the most suitable methods are penalized likelihood and Bayesian when a global clock assumption has been rejected, as these methods are more successful at finding optimal levels of smoothing to correct for rate heterogeneity, and are less sensitive to undersampling.

Africa↗

Genomic prediction and genome-wide association study for liver abscesses in crossbred beef cattle.

Liver abscesses are a concern in feedlot cattle, and little is known about the role of genetics in their development. This study aimed to estimate genetic parameters and to identify single-nucleotide polymorphisms (SNPs) associated with liver abscesses. Crossbred cattle representing 18 breeds in the U.S. Meat Animal Research Center Germplasm Evaluation Program were phenotyped for liver abscesses at slaughter (n&#x2005;=&#x2005;9,044). Seventeen percent of cattle had liver abscesses. These cattle had genotypes that were imputed to sequence variant genotypes. After filtering and quality control, 340,723 SNPs were used in the analysis. Liver abscess prevalence was modeled with a single-step genomic best linear unbiased prediction (ssGBLUP) threshold model using a Bayesian framework. The model included contemporary group (sex, treatment group, and slaughter date), additive genomic, and residual effects. Genomic heritability was 0.039 (95% highest posterior density&#x2005;=&#x2005;0.005, 0.081), which was very small. To assess prediction quality, a 5-fold random cross-validation structure was used. Method Linear Regression was used to assess accuracy, bias, and dispersion by comparing estimated breeding values (EBV) from full and reduced analyses. Cross-validation metrics showed EBV based on genotypes had 0.05 reliability (SD&#x2005;<&#x2005;0.01) with no bias relative to EBV based on genotypes and phenotypes. For the genome-wide association study, SNP effects were back calculated from the EBV solutions from ssGBLUP. No SNPs were associated with liver abscesses at a Benjamini-Hochberg adjusted 0.05 significance level. Although a large dataset was used, this result was because of the low genomic heritability and imprecise EBV used to calculate SNP effects. Based on these results, environmental factors contribute to most of the variation in liver abscesses. Genetic selection to reduce liver abscesses would be slow because of the low genomic heritability, measurement late in life, and inability to measure breeding animals. A faster approach would be finding additional environmental interventions that maintain animal performance.

Animals↗

A model to evaluate past exposure to 2,3,7,8-TCDD.

Data from several studies suggest that concentrations of dioxins rose in the environment from the 1930s to about the 1960s/70s and have been declining over the last decade or two. The most direct evidence of this trend comes from lake core sediments, which can be used to estimate past atmospheric depositions of dioxins. The primary source of human exposure to dioxins is through the food supply. The pathway relating atmospheric depositions to concentrations in food is quite complex, and accordingly, it is not known to what extent the trend in human exposure mirrors the trend in atmospheric depositions. This paper describes an attempt to statistically reconstruct the pattern of past human exposure to the most toxic dioxin congener, 2,3,7,8-TCDD (abbreviated TCDD), through use of a simple pharmacokinetic (PK) model which included a time-varying TCDD exposure dose. This PK model was fit to TCDD body burden data (i.e., TCDD concentrations in lipid) from five U.S. studies dating from 1972 to 1987 and covering a wide age range. A Bayesian statistical approach was used to fit TCDD exposure; model parameters other than exposure were all previously known or estimated from other data sources. The primary results of the analysis are as follows: (1) use of a time-varying exposure dose provided a far better fit to the TCDD body burden data than did using a dose that was constant over time; this is strong evidence that exposure to TCDD has, in fact, varied during the 20th century, (2) the year of peak TCDD exposure was estimated to be in the late 1960s, which coincides with peaks found in sediment core studies, (3) modeled average exposure doses during these peak years was estimated at 1.4-1.9 pg TCDD/kg-day, and (4) modeled exposure doses of TCDD for the late 1980s of less than 0.10 pg TCDD/kg-day correlated well with recent estimates of exposure doses around 0.17 pg TCDD/kg-day (recent estimates are based on food concentrations combined with food ingestion rates; food is thought to explain over 90% of total dioxin exposure). This paper describes these and other results, the goodness-of-fit between predicted and observed lipid TCDD concentrations, the modeled impact of breast feeding on lipid concentrations in young individuals, and sensitivity and uncertainty analyses.

Adolescent↗

Quantitative trait nucleotide analysis using Bayesian model selection.

Although much attention has been given to statistical genetic methods for the initial localization and fine mapping of quantitative trait loci (QTLs), little methodological work has been done to date on the problem of statistically identifying the most likely functional polymorphisms using sequence data. In this paper we provide a general statistical genetic framework, called Bayesian quantitative trait nucleotide (BQTN) analysis, for assessing the likely functional status of genetic variants. The approach requires the initial enumeration of all genetic variants in a set of resequenced individuals. These polymorphisms are then typed in a large number of individuals (potentially in families), and marker variation is related to quantitative phenotypic variation using Bayesian model selection and averaging. For each sequence variant a posterior probability of effect is obtained and can be used to prioritize additional molecular functional experiments. An example of this quantitative nucleotide analysis is provided using the GAW12 simulated data. The results show that the BQTN method may be useful for choosing the most likely functional variants within a gene (or set of genes). We also include instructions on how to use our computer program, SOLAR, for association analysis and BQTN analysis.

Bayes Theorem↗

Implementing the Bayesian paradigm: reporting research results over the World-Wide Web.

For decades, statisticians, philosophers, medical investigators and others interested in data analysis have argued that the Bayesian paradigm is the proper approach for reporting the results of scientific analyses for use by clients and readers. To date, the methods have been too complicated for non-statisticians to use. In this paper we argue that the World-Wide Web provides the perfect environment to put the Bayesian paradigm into practice: the likelihood function of the data is parsimoniously represented on the server side, the reader uses the client to represent her prior belief, and a downloaded program (a Java applet) performs the combination. In our approach, a different applet can be used for each likelihood function, prior belief can be assessed graphically, and calculation results can be reported in a variety of ways. We present a prototype implementation, BayesApplet, for two-arm clinical trials with normally-distributed outcomes, a prominent model for clinical trials. The primary implication of this work is that publishing medical research results on the Web can take a form beyond or different from that currently used on paper, and can have a profound impact on the publication and use of research results.

Bayes Theorem↗

Bayesian models of episodic evolution support a late precambrian explosive diversification of the Metazoa.

Multicellular animals, or Metazoa, appear in the fossil records between 575 and 509 million years ago (MYA). At odds with paleontological evidence, molecular estimates of basal metazoan divergences have been consistently older than 700 MYA. However, those date estimates were based on the molecular clock hypothesis, which is almost always violated. To relax this hypothesis, we have implemented a Bayesian approach to describe the change of evolutionary rate over time. Analysis of 22 genes from the nuclear and the mitochondrial genomes under the molecular clock assumption produced old date estimates, similar to those from previous studies. However, by allowing rates to vary in time and by taking small species-sampling fractions into account, we obtained much younger estimates, broadly consistent with the fossil records. In particular, the date of protostome-deuterostome divergence was on average 582 +/- 112 MYA. These results were found to be robust to specification of the model of rate change. The clock assumption thus had a dramatic effect on date estimation. However, our results appeared sensitive to the prior model of cladogenesis, although the oldest estimates (791 +/- 246 MYA) were obtained under a suboptimal model. Bayes posterior estimates of evolutionary rates indicated at least one major burst of molecular evolution at the end of the Precambrian when protostomes and deuterostomes diverged. We stress the importance of assumptions about rates on date estimation and suggest that the large discrepancies between the molecular and fossil dates of metazoan divergences might partly be due to biases in molecular date estimation.

Algorithms↗