Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

A bayesian approach to detect quantitative trait loci using Markov chain Monte Carlo.

Markov chain Monte Carlo (MCMC) techniques are applied to simultaneously identify multiple quantitative trait loci (QTL) and the magnitude of their effects. Using a Bayesian approach a multi-locus model is fit to quantitative trait and molecular marker data, instead of fitting one locus at a time. The phenotypic trait is modeled as a linear function of the additive and dominance effects of the unknown QTL genotypes. Inference summaries for the locations of the QTL and their effects are derived from the corresponding marginal posterior densities obtained by integrating the likelihood, rather than by optimizing the joint likelihood surface. This is done using MCMC by treating the unknown QTL, genotypes, and any missing marker genotypes, as augmented data and then by including these unknowns in the Markov chain cycle alone with the unknown parameters. Parameter estimates are obtained as means of the corresponding marginal posterior densities. High posterior density regions of the marginal densities are obtained as confidence regions. We examine flowering time data from double haploid progeny of Brassica napus to illustrate the proposed method.

Bayes Theorem↗

The psychology of good judgment: frequency formats and simple algorithms.

Mind and environment evolve in tandem--almost a platitude. Much of judgment and decision making research, however, has compared cognition to standard statistical models, rather than to how well it is adapted to its environment. The author argues two points. First, cognitive algorithms are tuned to certain information formats, most likely to those that humans have encountered during their evolutionary history. In particular, Bayesian computations are simpler when the information is in a frequency format than when it is in a probability format. The author investigates whether frequency formats can make physicians reason more often the Bayesian way. Second cognitive algorithms need to operate under constraints of limited time, knowledge, and computational power, and they need to exploit the structures of their environments. The author describes a fast and frugal algorithm. Take The Best, that violates standard principles of rational inference but can be as accurate as sophisticated "optimal" models for diagnostic inference.

Algorithms↗

Bayesian learning of sparse gene regulatory networks.

Differential equations (DEs) have been the most widespread formalism for gene regulatory network (GRN) modeling, as they offer natural interpretation of biological processes, easy elucidation of gene relationships, and the capability of using efficient parameter estimation methods. However, an important limitation of DEs is their requirement of O(d(2)) parameters where d is the number of genes modeled, which often causes over-parameterization for large d, leading to the over-fitting of data and dense parameter sets that are hard to interpret. This paper presents the first effort to address the over-parameterization problem by applying the sparse Bayesian learning (SBL) method to sparsify the GRN model of DEs. SBL operates on the parsimony principle, with the objective to reduce the number of effective parameters by driving the redundant parameters to zero. The resulting sparse parameter set offers three important advantages for GRN inference: first, the inferred GRNs are more plausible, since the biological counterparts are known to be sparse; second, gene relationships can be more easily elucidated from sparse sets than from dense sets; and third, the solutions become more optimal and consistent, due to the reduction in the volume of solution space. Experiments are conducted on the yeast Saccharomyces cerevisiae time-series gene expression data, in which known regulatory events related to the cell cycle G1/S phase are reliably reproduced.

Bayes Theorem↗

A molecular phylogenetic analysis of strombid gastropod morphological diversity.

The shells of strombid gastropods show a wide variety of forms, ranging from small and fusiform to large and elaborately ornamented with a strongly flared outer lip. Here, we present the first species-level molecular phylogeny for strombids and use the resulting phylogenetic framework to explore relationships between species richness and morphological diversity. We use portions of one nuclear (325 bp of histone H3) and one mitochondrial (640 bp of cytochrome oxidase I, COI) gene to infer relationships within the two most species-rich genera in the Strombidae: Strombus and Lambis. We include 32 species of Strombus, representing 10 of 11 extant subgenera, and 3 of the 9 species of Lambis, representing 2 of 3 extant subgenera. Maximum likelihood and Bayesian analyses of COI and of H3 and COI combined suggest Lambis is nested within a paraphyletic Strombus. Eastern Pacific and western Atlantic species of Strombus form a relatively recent monophyletic radiation within an older, paraphyletic Indo-West Pacific grade. Morphological diversity of subclades scales positively with species richness but does not show evidence of strong phylogenetic constraints.

Animals↗

Evidence of genetic distinction and long-term population decline in wolves (Canis lupus) in the Italian Apennines.

Historical information suggests the occurrence of an extensive human-caused contraction in the distribution range of wolves (Canis lupus) during the last few centuries in Europe. Wolves disappeared from the Alps in the 1920s, and thereafter continued to decline in peninsular Italy until the 1970s, when approximately 100 individuals survived, isolated in the central Apennines. In this study we performed a coalescent analysis of multilocus DNA markers to infer patterns and timing of historical population changes in wolves surviving in the Apennines. This population showed a unique mitochondrial DNA control-region haplotype, the absence of private alleles and lower heterozygosity at microsatellite loci, as compared to other wolf populations. Multivariate, clustering and Bayesian assignment procedures consistently assigned all the wolf genotypes sampled in Italy to a single group, supporting their genetic distinction. Bottleneck tests showed evidences of population decline in the Italian wolves, but not in other populations. Results of a Bayesian coalescent model indicate that wolves in Italy underwent a 100- to 1000-fold population contraction over the past 2000-10,000 years. The population decline was stronger and longer in peninsular Italy than elsewhere in Europe, suggesting that wolves have apparently been genetically isolated for thousands of generations south of the Alps. Ice caps covering the Alps at the Last Glacial Maximum (c. 18,000 years before present), and the wide expansion of the Po River, which cut the alluvial plains throughout the Holocene, might have provided effective geographical barriers to wolf dispersal. More recently, the admixture of Alpine and Apennine wolf populations could have been prevented by deforestation, which was already widespread in the fifteenth century in northern Italy. This study suggests that, despite the high potential rates of dispersal and gene flow, local wolf populations may not have mixed for long periods of time.

Animals↗

A molecular phylogeny of the genus Echinococcus inferred from complete mitochondrial genomes.

Taxonomic revision by molecular phylogeny is needed to categorize members of the genus Echinococcus (Cestoda: Taeniidae). We have reconstructed the phylogenetic relationships of E. oligarthrus, E. vogeli, E. multilocularis, E. shiquicus, E. equinus, E. ortleppi, E. granulosus sensu stricto and 3 genotypes of E. granulosus sensu lato (G6, G7 and G8) from their complete mitochondrial genomes. Maximum likelihood and partitioned Bayesian analyses using concatenated data sets of nucleotide and amino acid sequences depicted phylogenetic trees with the same topology. The 3 E. granulosus genotypes corresponding to the camel, pig, and cervid strains were monophyletic, and their high level of genetic similarity supported taxonomic species unification of these genotypes into E. canadensis. Sister species relationships were confirmed between E. ortleppi and E. canadensis, and between E. multilocularis and E. shiquicus, regardless of the analytical approach employed. The basal positions of the phylogenetic tree were occupied by the neotropical endemic species, E. oligarthrus and E. vogeli, whose definitive hosts are derived from carnivores that immigrated from North America after the formation of the Panamanian land bridge. Host-parasite co-evolution comparisons suggest that the ancestral homeland of Echinococcus was North America or Asia, depending on whether the ancestral definitive hosts were canids or felids.

Animals↗

Mating patterns, relatedness and the basis of natal philopatry in the brown long-eared bat, Plecotus auritus.

The brown long-eared bat, Plecotus auritus, is unusual among temperate zone bats in that summer maternity colonies are composed of adult males and females, with both sexes displaying natal philopatry and long-term association with a colony. Here, we describe the use of microsatellite analysis to investigate colony relatedness and mating patterns, with the aim of identifying the evolutionary determinants of social organization in P. auritus. Mean colony relatedness was found to be low (R=0.033 +/- 0.002), with pairwise estimates of R within colonies ranging from -0.4 to 0.9. The proportion of young fathered by males in their own colony was investigated using a Bayesian approach, incorporating parameters detailing the number of untyped individuals. This analysis revealed that most offspring were fathered by males originating from a different colony to their own. In addition, we determined that the number of paternal half-sibs among cohorts of young was low, inferring little or no skew in male reproductive success. The results of this study suggest that kin selection cannot account for colony stability and natal philopatry in P. auritus, which may instead be explained by advantages accrued through the use of familiar and successful roost sites, and through long-term associations with conspecifics. Moreover, because the underlying causes of male natal dispersal in mammals, such as risk of inbreeding or competition for mates, appear to be avoided via extra-colony copulation and low male reproductive skew, both P. auritus males and females are able to benefit from long-term association with the natal colony.

Animals↗

The F7 gene and clotting factor VII levels: dissection of a human quantitative trait locus.

Localization of human quantitative trait loci (QTLs) is now routine. However, identifying their functional DNA variants is still a formidable challenge. We present a complete dissection of a human QTL using novel statistical techniques to infer the most likely functional polymorphisms of a QTL that influence plasma levels of clotting factor VII (FVII), a risk factor for cardiovascular disease. Resequencing of 15 kb in and around the F7 gene identified 49 polymorphisms, which were then genotyped in 398 people. Using a Bayesian quantitative trait nucleotide (BQTN) method, we identified four to seven functional variants that completely account for this QTL. These variants include both rare coding variants and more common, potentially regulatory polymorphisms in intronic and promoter regions.

Bayes Theorem↗

Inferring subnetworks from perturbed expression profiles.

Genome-wide expression profiles of genetic mutants provide a wide variety of measurements of cellular responses to perturbations. Typical analysis of such data identifies genes affected by perturbation and uses clustering to group genes of similar function. In this paper we discover a finer structure of interactions between genes, such as causality, mediation, activation, and inhibition by using a Bayesian network framework. We extend this framework to correctly handle perturbations, and to identify significant subnetworks of interacting genes. We apply this method to expression data of S. cerevisiae mutants and uncover a variety of structured metabolic, signaling and regulatory pathways.

Bayes Theorem↗

Phylogenetic relationships and evolutionary history of snake-eyed skink Ablepharus kitaibelii (Sauria: Scincidae).

Sequence data derived from two mitochondrial markers, 16S rRNA and cytochrome b genes, were used to infer the phylogenetic relationships of 38 populations of the snake-eyed skinks of the genus Ablepharus with emphasis on A. kitaibelii from Greece and Turkey. The partition-homogeneity tests indicated that the combined data set was homogeneous, and maximum-parsimony, maximum-likelihood, and Bayesian analyses produced topologically identical trees that revealed a well-resolved phylogeny. All species except A. kitaibelii form monophyletic units. The latter species appears paraphyletic with respect to A. budaki and A. chernovi with populations clustering into two distinct clades. A. chernovi and A. budaki, which have recently been raised to species status, were confirmed as genetically distinct forms. We used sequence divergence and paleogeographic history of the Aegean region to reconstruct a biogeographic evolutionary scenario for A. kitaibelii.

Animals↗

Phylogeny of Rhaponticum (Asteraceae, Cardueae-Centaureinae) and related genera inferred from nuclear and chloroplast DNA sequence data: taxonomic and biogeographic implications.

BACKGROUND AND AIMS: The precise generic delimitation of the Rhaponticum group is not totally resolved. The lack of knowledge of the relationships between the basal genera of Centaureinae could imply that genera whose position is as yet unresolved could belong to the Rhaponticum group. On the other hand, the affinities among the genera that are considered as members of this group are not well known. The aim of the study is to contribute to the phylogenetic and generic delineation of the Rhaponticum group on the basis of molecular data. METHODS: Parsimony and Bayesian analyses of the combined sequences of one plastid (trnL-trnF) and two nuclear (ITS region and ETS) molecular markers were carried out. The results of these analyses are discussed in the light of the biogeographic history. KEY RESULTS: The Rhaponticum group appears as monophyletic, and closely related to the genus Klasea. The results confirm the preliminary generic delimitation of the Rhaponticum group, with the new incorporation of the genus Centaurothamnus. Ochrocephala is supported as a separate genus from Rhaponticum and, contrary to this, Acroptilon and Leuzea appear as merged into the genus Rhaponticum. Several nomenclatural rearrangements are made in Klasea and Rhaponticum. CONCLUSIONS: The new molecular evidence is consistent with the morphological and karyological data, and suggests particularly coherent biogeographic routes of migration and speciation processes for the genus Rhaponticum. The biogeographic inference proposes a Near East and/or Caucasian origin for the genus. Furthermore, representatives of Rhaponticum could have reached Europe in two different ways: (1) expansion across central Asia to eastern Europe, and (2) expansion through the Near East, North Africa and then to the Iberian Peninsula and the Alps.

Asteraceae↗

Phylogenetic relationships of the macaques (Cercopithecidae: Macaca), inferred from mitochondrial DNA sequences.

To study the phylogenetic relationships of the macaques, five gene fragments were sequenced from 40 individuals of eight species: Macaca mulatta, M. cyclopis, M. fascicularis, M. arctoides, M. assamensis, M. thibetana, M. silenus, and M. leonina. In addition, sequences of M. sylvanus were obtained from Genbank. A baboon was used as the outgroup. The phylogenetic trees were constructed using maximum-parsimony and Bayesian methods. Because five gene fragments were from the mitochondrial genome and were inherited as a single entity without recombination, we combined the five genes into a single analysis. The parsimony bootstrap proportions we obtained were higher than those from earlier studies based on the combined mtDNA dataset. Excluding M. arctoides, our results are generally consistent with the classification of Delson (1980). Our phylogenetic analyses agree with earlier studies suggesting that the mitochondrial lineages of M. arctoides share a close evolutionary relationship with the mitochondrial lineages of the fascicularis group of macaques (and M. fascicularis, specifically). M. mulatta (with respect to M. cyclopis), M. assamensis assamensis (with respect to M. thibetana), and M. leonina (with respect to M. silenus) are paraphyletic based on our analysis of mitochondrial genes.

Animals↗

Adaptive evolution in the Arabidopsis MADS-box gene family inferred from its complete resolved phylogeny.

Gene duplication is a substrate of evolution. However, the relative importance of positive selection versus relaxation of constraints in the functional divergence of gene copies is still under debate. Plant MADS-box genes encode transcriptional regulators key in various aspects of development and have undergone extensive duplications to form a large family. We recovered 104 MADS sequences from the Arabidopsis genome. Bayesian phylogenetic trees recover type II lineage as a monophyletic group and resolve a branching sequence of monophyletic groups within this lineage. The type I lineage is comprised of several divergent groups. However, contrasting gene structure and patterns of chromosomal distribution between type I and II sequences suggest that they had different evolutionary histories and support the placement of the root of the gene family between these two groups. Site-specific and site-branch analyses of positive Darwinian selection (PDS) suggest that different selection regimes could have affected the evolution of these lineages. We found evidence for PDS along the branch leading to flowering time genes that have a direct impact on plant fitness. Sites with high probabilities of having been under PDS were found in the MADS and K domains, suggesting that these played important roles in the acquisition of novel functions during MADS-box diversification. Detected sites are targets for further experimental analyses. We argue that adaptive changes in MADS-domain protein sequences have been important for their functional divergence, suggesting that changes within coding regions of transcriptional regulators have influenced phenotypic evolution of plants.

Arabidopsis↗

PhyClone: accurate Bayesian reconstruction of cancer phylogenies from bulk sequencing.

MOTIVATION: Cancer is driven by somatic mutations that result in the expansion of genomically distinct sub-populations of cells called clones. Identifying the clonal composition of tumours and understanding the evolutionary relationships between clones is a crucial task in cancer genomics. Bulk DNA sequencing is commonly used for studying the clonal composition of tumours, but it is challenging to infer the genetic relationship between different clones due to the mixture of different cell populations. RESULTS: In this work, we introduce a new probabilistic model called PhyClone that can infer clonal phylogenies from bulk-sequencing data. We demonstrate the performance of PhyClone on simulated and real-world datasets and show that it outperforms previous methods in terms of accuracy and sample scalability. AVAILABILITY AND IMPLEMENTATION: Source code is available on Github at: https://github.com/Roth-Lab/PhyClone under the GPL v3.0 license.

Neoplasms↗

Using linked markers to infer the age of a mutation.

Advances in sequencing and genotyping technologies over the last decade have enabled geneticists to easily characterize genetic variation at the nucleotide level. Hundreds of genes harboring mutations associated with genetic disease have now been identified by positional cloning. Using variation at closely linked genetic markers, it is possible to predict the times in the past at which particular mutations arose. Such studies suggest that many of the rare mutations underlying human genetic disorders are relatively young. Studies of variation at genetic markers linked to particular mutations can provide insights into human geographic history, and historical patterns of natural selection and disease, that are not available from other sources. We review two approaches for estimating allele age using variation at linked genetic markers. A phylogenetic approach aims to reconstruct the gene tree underlying a sample of chromosomes carrying a particular mutation, obtaining a "direct" estimate of allele age from the age of the root of this tree. A population genetic approach relies on models of demography, mutation, and/or recombination to estimate allele age without explicitly reconstructing the gene tree. Phylogenetic methods are best suited for studies of ancient mutations, while population genetic methods are better suited for studies of recent mutations. Methods that rely on recombination to infer the ages of alleles can be fine-tuned by choosing linked markers at optimal map distances to maximize the information available about allele age. A limitation of methods that rely on recombination is the frequent lack of a fine-scale linkage map. Maximum likelihood and Bayesian methods for estimating allele age that rely on intensive numerical computation are described, as well as "composite" likelihood and moment-based methods that lead to simple estimators. The former provide more accurate estimates (particularly for large samples of chromosomes) and should be employed if computationally practical.

Alleles↗

Phylogenomics and molecular evolution of polyomaviruses.

We provide in this chapter an overview of the basic steps to reconstruct evolutionary relationships through standard phylogeny estimation approaches as well as network approaches for sequences more closely related. We discuss the importance of sequence alignment, selecting models of evolution, and confidence assessment in phylogenetic inference. We also introduce the reader to a variety of software packages used for such studies. Finally, we demonstrate these approaches throughout using a data set of 33 whole genomes of polyomaviruses. A robust phylogeny of these genomes is estimated and phylogenetic relationships among the polyomaviruses determined using Bayesian and maximum likelihood approaches. Furthermore, population samples of SV40 are used to demonstrate the utility of network approaches for closely related sequences. The phylogenetic analysis suggested a close relationship among the BK viruses, JC viruses, and SV40 with a more distant association with mouse polyomavirus, monkey polymavirus (LPV) and then avian polyomavirus (BFDV).

Computational Biology↗

Chromosome painting and molecular dating indicate a low rate of chromosomal evolution in golden moles (Mammalia, Chrysochloridae).

Golden moles (Chrysochloridae) are poorly known subterranean mammals endemic to Southern Africa that are part of the superordinal clade Afrotheria. Using G-banding and chromosome painting we provide a comprehensive comparison of the karyotypes of five species representing five of the nine recognized genera: Amblysomus hottentotus, Chrysochloris asiatica, Chrysospalax trevelyani, Cryptochloris zyli and Eremitalpa granti. The species are karyotypically highly conserved. In total, only four changes were detected among them. Eremitalpa granti has the most derived karyotype with 2n = 26 and differs from the remaining species (all of whom have 2n = 30) by one centric and one telomere:telomere fusion. In addition, two intrachromosomal rearrangements were detected in A. hottentotus. The painting probes also suggest the presence of a unique satellite DNA family located on chromosomes 11 and 12 of both C. asiatica and C. zyli. This represents a synapomorphy linking these two sympatric species as sister taxa. A molecular clock was calibrated adopting a relaxed Bayesian approach for multigene data sets comprising publicly available sequences derived from five gene fragments representative of three golden moles and 39 other eutherian species. The data suggest that golden moles diverged from a common ancestor approximately 28.5 mya (95% credibility interval = 21.5-36.5 mya). Based on an inferred chrysochlorid ancestral karyotype of 2n = 30, the estimated rate of 0.7 rearrangements per 10 my (95% Credibility Interval = 0.54-0.93) differs from the 'default rate' of mammalian chromosomal evolution which has been estimated at one change per 10 million years, thus placing the Chrysochloridae among the slower-evolving chromosomal lineages thus far recorded.

Animals↗

A simple statistical method for estimating type-II (cluster-specific) functional divergence of protein sequences.

Predicting functional amino acid residues in silico is important for comparative genomics. In this paper, we focus on the issue of how to statistically identify cluster-specific amino acid residues that are related to the functional divergence after gene duplication. We approach this problem using a framework based on site-specific shift of amino acid property (type-II functional divergence), as opposed to site-specific shift of evolutionary rate (type-I functional divergence). An efficient statistical procedure is implemented to facilitate the development of phylogenomic database for cluster-specific residues of large-scale protein families. Our method has the following features: 1) statistical testing of the type-II functional divergence and 2) the site-specific Bayesian profile to measure how amino acid residues contribute to type-II (cluster-specific) functional divergence. Consequently, one may obtain the posterior probability for "functional" cluster-specific residues. Case studies are presented and indicate that radical cluster-specific residues are responsible for most of inferred type-II functional divergence, whereas conserved cluster-specific residues appear less than even those imperfect radical cluster-specific residues to this type of functional divergence.

Amino Acid Sequence↗