Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Detecting Interspecific Positive Selection Using Convolutional Neural Networks.

Traditional statistical methods using maximum likelihood and Bayesian inference can detect positive selection from an interspecific phylogeny and a codon sequence alignment based on model assumptions, but they are prone to false positives due to alignment errors and can lack power. These problems are particularly pronounced when faced with high levels of indels and divergence. To address these issues, we trained and tested convolutional neural network models on simulated data and achieved higher accuracy in detecting selection across a specific range of phylogenetic scenarios and evolutionary modes. This advantage is particularly evident when performing inference on noisy data prone to misalignments. Our method shows some ability to account for these errors, where most statistical frameworks fail to do so in a tractable manner. We explore the generalizability of our convolutional neural network models to unseen evolutionary scenarios and identify future avenues to achieve broader utility. Once trained, our convolutional neural network model is faster at test time, making it a scalable alternative to traditional statistical methods for large-scale, multigene analyses. In addition to binary classification (inference of the presence or absence of positive selection during the evolution of the sequences), we use saliency maps to understand what the model learns and observe how this could be leveraged for sitewise inference of positive selection.

Neural Networks, Computer↗

The epidemiology and iatrogenic transmission of hepatitis C virus in Egypt: a Bayesian coalescent approach.

Hepatitis C virus (HCV) is a leading cause of liver cancer and cirrhosis, and Egypt has possibly the highest HCV prevalence worldwide. In this article we use a newly developed Bayesian inference framework to estimate the transmission dynamics of HCV in Egypt from sampled viral gene sequences, and to predict the public health impact of the virus. Our results indicate that the effective number of HCV infections in Egypt underwent rapid exponential growth between 1930 and 1955. The timing and speed of this spread provides quantitative genetic evidence that the Egyptian HCV epidemic was initiated and propagated by extensive antischistosomiasis injection campaigns. Although our results show that HCV transmission has since decreased, we conclude that HCV is likely to remain prevalent in Egypt for several decades. Our combined population genetic and epidemiological analysis provides detailed estimates of historical changes in Egyptian HCV prevalence. Because our results are consistent with a demographic scenario specified a priori, they also provide an objective test of inference methods based on the coalescent process.

Bayes Theorem↗

A molecular timeline for the origin of photosynthetic eukaryotes.

The appearance of photosynthetic eukaryotes (algae and plants) dramatically altered the Earth's ecosystem, making possible all vertebrate life on land, including humans. Dating algal origin is, however, frustrated by a meager fossil record. We generated a plastid multi-gene phylogeny with Bayesian inference and then used maximum likelihood molecular clock methods to estimate algal divergence times. The plastid tree was used as a surrogate for algal host evolution because of recent phylogenetic evidence supporting the vertical ancestry of the plastid in the red, green, and glaucophyte algae. Nodes in the plastid tree were constrained with six reliable fossil dates and a maximum age of 3,500 MYA based on the earliest known eubacterial fossil. Our analyses support an ancient (late Paleoproterozoic) origin of photosynthetic eukaryotes with the primary endosymbiosis that gave rise to the first alga having occurred after the split of the Plantae (i.e., red, green, and glaucophyte algae plus land plants) from the opisthokonts sometime before 1,558 MYA. The split of the red and green algae is calculated to have occurred about 1,500 MYA, and the putative single red algal secondary endosymbiosis that gave rise to the plastid in the cryptophyte, haptophyte, and stramenopile algae (chromists) occurred about 1,300 MYA. These dates, which are consistent with fossil evidence for putative marine algae (i.e., acritarchs) from the early Mesoproterozoic (1,500 MYA) and with a major eukaryotic diversification in the very late Mesoproterozoic and Neoproterozoic, provide a molecular timeline for understanding algal evolution.

Bayes Theorem↗

The chloroplast genome of Phalaenopsis aphrodite (Orchidaceae): comparative analysis of evolutionary rate with that of grasses and its phylogenetic implications.

Whether the Amborella/Amborella-Nymphaeales or the grass lineage diverged first within the angiosperms has recently been debated. Central to this issue has been focused on the artifacts that might result from sampling only grasses within the monocots. We therefore sequenced the entire chloroplast genome (cpDNA) of Phalaenopsis aphrodite, Taiwan moth orchid. The cpDNA is a circular molecule of 148,964 bp with a comparatively short single-copy region (11,543 bp) due to the unusual loss and truncation/scattered deletion of certain ndh subunits. An open reading frame, orf91, located in the complementary strand of the rrn23 was reported for the first time. A comparison of nucleotide substitutions between P. aphrodite and the grasses indicates that only the plastid expression genes have a strong positive correlation between nonsynonymous (Ka) and synonymous (Ks) substitutions per site, providing evidence for a generation time effect, mainly across these genes. Among the intron-containing protein-coding genes of the sampled monocots, the Ks of the genes are significantly correlated to transitional substitutions of their introns. We compiled a concatenated 61 protein-coding gene alignment for the available 20 cpDNAs of vascular plants and analyzed the data set using Bayesian inference, maximum parsimony, and neighbor-joining (NJ) methods. The analyses yielded robust support for the Amborella/Amborella-Nymphaeales-basal hypothesis and for the orchid and grasses together being a monophyletic group nested within the remaining angiosperms. However, the NJ analysis using Ka, the first two codon positions, or amino acid sequences, respectively, supports the monocots-basal hypothesis. We demonstrated that these conflicting angiosperm phylogenies are most probably linked to the transitional sites at all codon positions, especially at the third one where the strong base-composition bias and saturation effect take place.

DNA, Chloroplast↗

Evolution of glutamine synthetase in heterokonts: evidence for endosymbiotic gene transfer and the early evolution of photosynthesis.

Although the endosymbiotic evolution of chloroplasts through primary and secondary associations is well established, the evolutionary timing and stability of the secondary endosymbiotic events is less well resolved. Heterokonts include both photosynthetic and nonphotosynthetic members and the nonphotosynthetic lineages branch basally in phylogenetic reconstructions. Molecular and morphological data indicate that heterokont chloroplasts evolved via a secondary endosymbiosis, involving a heterotrophic host cell and a photosynthetic ancestor of the red algae and this endosymbiotic event may have preceded the divergence of heterokonts and alveolates. If photosynthesis evolved early in this lineage, nuclear genomes of the nonphotosynthetic groups may contain genes that are not essential to photosynthesis but were derived from the endosymbiont genome through gene transfer. These genes offer the potential to trace the evolutionary history of chloroplast gains and losses within these lineages. Glutamine synthetase (GS) is essential for ammonium assimilation and glutamine biosynthesis in all organisms. Three paralogous gene families (GSI, GSII, and GSIII) have been identified and are broadly distributed among prokaryotic and eukaryotic lineages. In diatoms (Heterokonta), the nuclear-encoded chloroplast and cytosolic-localized GS isoforms are encoded by members of the GSII and GSIII family, respectively. Here, we explore the evolutionary history of GSII in both photosynthetic and nonphotosynthetic heterokonts, red algae, and other eukaryotes. GSII cDNA sequences were obtained from two species of oomycetes by polymerase chain reaction amplification. Additional GSII sequences from eukaryotes and bacteria were obtained from publicly available databases and genome projects. Bayesian inference and maximum likelihood phylogenetic analyses of GSII provided strong support for the monophyly of heterokonts, rhodophytes, chlorophytes, and plants and strong to moderate support for the Opisthokonts. Although the phylogeny is reflective of the unikont/bikont division of eukaryotes, we propose based on the robustness of the phylogenetic analyses that the heterokont GSII gene evolved via endosymbiotic gene transfer from the nucleus of the red-algal endosymbiont to the nucleus of the host. The lack of GSIII sequences in the oomycetes examined here further suggests that the GSIII gene that functions in the cytosol of photosynthetic heterokonts was replaced by the endosymbiont-derived GSII gene.

Amino Acid Sequence↗

Dependence among sites in RNA evolution.

Although probabilistic models of genotype (e.g., DNA sequence) evolution have been greatly elaborated, less attention has been paid to the effect of phenotype on the evolution of the genotype. Here we propose an evolutionary model and a Bayesian inference procedure that are aimed at filling this gap. In the model, RNA secondary structure links genotype and phenotype by treating the approximate free energy of a sequence folded into a secondary structure as a surrogate for fitness. The underlying idea is that a nucleotide substitution resulting in a more stable secondary structure should have a higher rate than a substitution that yields a less stable secondary structure. This free energy approach incorporates evolutionary dependencies among sequence positions beyond those that are reflected simply by jointly modeling change at paired positions in an RNA helix. Although there is not a formal requirement with this approach that secondary structure be known and nearly invariant over evolutionary time, computational considerations make these assumptions attractive and they have been adopted in a software program that permits statistical analysis of multiple homologous sequences that are related via a known phylogenetic tree topology. Analyses of 5S ribosomal RNA sequences are presented to illustrate and quantify the strong impact that RNA secondary structure has on substitution rates. Analyses on simulated sequences show that the new inference procedure has reasonable statistical properties. Potential applications of this procedure, including improved ancestral sequence inference and location of functionally interesting sites, are discussed.

Animals↗

The ProDom database of protein domain families: more emphasis on 3D.

ProDom is a comprehensive database of protein domain families generated from the global comparison of all available protein sequences. Recent improvements include the use of three-dimensional (3D) information from the SCOP database; a completely redesigned web interface (http://www.toulouse.inra.fr/prodom.html); visualization of ProDom domains on 3D structures; coupling of ProDom analysis with the Geno3D homology modelling server; Bayesian inference of evolutionary scenarios for ProDom families. In addition, we have developed ProDom-SG, a ProDom-based server dedicated to the selection of candidate proteins for structural genomics.

Computer Graphics↗

Prediction of functional modules based on comparative genome analysis and Gene Ontology application.

We present a computational method for the prediction of functional modules encoded in microbial genomes. In this work, we have also developed a formal measure to quantify the degree of consistency between the predicted and the known modules, and have carried out statistical significance analysis of consistency measures. We first evaluate the functional relationship between two genes from three different perspectives--phylogenetic profile analysis, gene neighborhood analysis and Gene Ontology assignments. We then combine the three different sources of information in the framework of Bayesian inference, and we use the combined information to measure the strength of gene functional relationship. Finally, we apply a threshold-based method to predict functional modules. By applying this method to Escherichia coli K12, we have predicted 185 functional modules. Our predictions are highly consistent with the previously known functional modules in E.coli. The application results have demonstrated that our approach is highly promising for the prediction of functional modules encoded in a microbial genome.

Bayes Theorem↗

Hypothesis: bacterial clamp loader ATPase activation through DNA-dependent repositioning of the catalytic base and of a trans-acting catalytic threonine.

The prokaryotic DNA polymerase III clamp loader complex loads the beta clamp onto DNA to link the replication complex to DNA during processive synthesis and unloads it again once synthesis is complete. This minimal complex consists of one delta, one delta' and three gamma subunits, all of which possess an AAA+ module--though only the gamma subunit exhibits ATPase activity. Here clues to underlying clamp loader mechanisms are obtained through Bayesian inference of various categories of selective constraints imposed on the gamma and delta' subunits. It is proposed that a conserved histidine is ionized via electron transfer involving structurally adjacent residues within the sensor 1 region of gamma's AAA+ module. The resultant positive charge on this histidine inhibits ATPase activity by drawing the negatively charged catalytic base away from the active site. It is also proposed that this arrangement is disrupted upon interaction of DNA with basic residues in gamma implicated previously in DNA binding, regarding which a lysine that is near the sensor 1 region and that is highly conserved both in bacterial and in eukaryotic clamp loader ATPases appears to play a critical role. gamma ATPases also appear to utilize a trans-acting threonine that is donated by helix 6 of an adjacent gamma or delta' subunit and that assists in the activation of a water molecule for nucleophilic attack on the gamma phosphorous atom of ATP. As eukaryotic and archaeal clamp loaders lack most of these key residues, it appears that eubacteria utilize a fundamentally different mechanism for clamp loader activation than do these other organisms.

Adenosine Triphosphatases↗

Accounting for pregnancy dependence in epidemiologic studies of reproductive outcomes.

The aim of this paper is to evaluate the contribution of hierarchical mixed models to the analysis of epidemiologic studies of environmental exposure and reproductive outcomes. We have re-analyzed, with a logistic-normal mixed model, four studies investigating the relation between the frequency of spontaneous abortions and paternal or maternal environmental exposures. The data include multiple pregnancies for some women. The fitted models allow for between-woman variation of the propensity for spontaneous abortion, by including a random intercept in the logistic model to adjust for within-woman correlations on pregnancy outcomes. We have discussed and implemented two estimation methods, maximum likelihood and Bayesian inference. We found similar values in the various epidemiologic studies of the between-woman variance of the intrinsic risk of spontaneous abortion. The size of this variance corresponds to a substantial variability in risk between women. Indeed, the risk of spontaneous abortion calculated for "nonexposed" pregnancies, that is, with mother's age, birth order, tobacco consumption, and maternal environmental exposure equal to the referent class, can vary, according to this model, from 2% to 17%.

Abortion, Spontaneous↗

Systems biology for cancer.

PURPOSE OF REVIEW: Significant insight can be gained into complex biologic mechanisms of cancer via a combined computational and experimental systems biology approach. This review highlights some of the major systems biology efforts that were applied to cancer in the past year. RECENT FINDINGS: Two main approaches to computational systems biology are discussed: mechanistic dynamical simulations and inferential data mining. Significant developments have occurred in both areas. For example, mechanistic simulations of the EGFR pathway are promoting understanding of cancer, and Bayesian inference approaches allow for the reconstruction of regulatory networks. In addition, the article reports on advancements in experimental systems biology for determining protein-protein interactions and quantifying protein expression to generate the necessary data for computational modeling and inferential data mining. Emerging approaches will further improve the ability to bridge the gap between in vitro systems and in vivo human biology. Technologies paving the way include in vitro models that better reflect in vivo tumors, microfabricated devices of human physiology, and improved animal models. SUMMARY: An important challenge facing the field is how better to translate in vitro discoveries to the clinic. Computational systems biology approaches that use omic data to predict biology along with novel experimental systems that better represent human in vivo biology will prove useful in bridging this gap. Although still early, the potential application of systems biology and the future evolution of the field will significantly affect understanding of cancer disease mechanisms and the ability to devise effective therapeutics.

Computational Biology↗

Multiple gene evidence for expansion of extant penguins out of Antarctica due to global cooling.

Classic problems in historical biogeography are where did penguins originate, and why are such mobile birds restricted to the Southern Hemisphere? Competing hypotheses posit they arose in tropical-warm temperate waters, species-diverse cool temperate regions, or in Gondwanaland approximately 100 mya when it was further north. To test these hypotheses we constructed a strongly supported phylogeny of extant penguins from 5851 bp of mitochondrial and nuclear DNA. Using Bayesian inference of ancestral areas we show that an Antarctic origin of extant taxa is highly likely, and that more derived taxa occur in lower latitudes. Molecular dating estimated penguins originated about 71 million years ago in Gondwanaland when it was further south and cooler. Moreover, extant taxa are inferred to have originated in the Eocene, coincident with the extinction of the larger-bodied fossil taxa as global climate cooled. We hypothesize that, as Antarctica became ice-encrusted, modern penguins expanded via the circumpolar current to oceanic islands within the Antarctic Convergence, and later to the southern continents. Thus, global cooling has had a major impact on penguin evolution, as it has on vertebrates generally. Penguins only reached cooler tropical waters in the Galapagos about 4 mya, and have not crossed the equatorial thermal barrier.

Animal Migration↗

Colour-assortative mating among populations of Tropheus moorii, a cichlid fish from Lake Tanganyika, East Africa.

The species flocks of cichlid fishes in the East African Lakes Tanganyika, Malawi and Victoria are prime examples of adaptive radiation and explosive speciation. Several hundreds of endemic species have evolved in each of the lakes over the past several thousands to a few millions years. Sexual selection via colour-assortative mating has often been proposed as a probable causal factor for initiating and maintaining reproductive isolation. Here, we report the consequences of human-mediated admixis among differentially coloured populations of the endemic cichlid fish Tropheus moorii from several localities that have accidentally been put in sympatry in a small harbour bay in the very south of Lake Tanganyika. We analysed the phenotypes (coloration) and genotypes (mitochondrial control region and five microsatellite loci) of almost 500 individuals, sampled over 3 consecutive years. Maximum-likelihood-based parenthood analyses and Bayesian inference of population structure revealed that significantly more juveniles are the product of within-colour-morph matings than could be expected under the assumption of random mating. Our results clearly indicate a marked degree of assortative mating with respect to the different colour morphs. Therefore, we postulate that sexual selection based on social interactions and female mate choice has played an important role in the formation and maintenance of the different colour morphs in Tropheus, and is probably common in other maternally mouthbrooding cichlids as well.

Alleles↗

Monte Carlo calibration of avalanches described as Coulomb fluid flows.

The idea that snow avalanches might behave as granular flows, and thus be described as Coulomb fluid flows, came up very early in the scientific study of avalanches, but it is not until recently that field evidence has been provided that demonstrates the reliability of this idea. This paper aims to specify the bulk frictional behaviour of snow avalanches by seeking a universal friction law. Since the bulk friction coefficient cannot be measured directly in the field, the friction coefficient must be calibrated by adjusting the model outputs to closely match the recorded data. Field data are readily available but are of poor quality and accuracy. We used Bayesian inference techniques to specify the model uncertainty relative to data uncertainty and to robustly and efficiently solve the inverse problem. A sample of 173 events taken from seven paths in the French Alps was used. The first analysis showed that the friction coefficient behaved as a random variable with a smooth and bell-shaped empirical distribution function. Evidence was provided that the friction coefficient varied with the avalanche volume, but any attempt to adjust a one-to-one relationship relating friction to volume produced residual errors that could be as large as three times the maximum uncertainty of field data. A tentative universal friction law is proposed: the friction coefficient is a random variable, the distribution of which can be approximated by a normal distribution with a volume-dependent mean.

Complex Mixtures↗

Repeat-type distribution in trnL intron does not correspond with species phylogeny: comparison of the genetic markers 16S rRNA and trnL intron in heterocystous cyanobacteria.

tRNA(Leu) UAA (trnL) intron sequences are used as genetic markers for differentiating cyanobacteria and for constructing phylogenies, since the introns are thought to be more variable among close relatives than is the 16S rRNA gene, the conventional phylogenetic marker. The evolution of trnL intron sequences and their utility as a phylogenetic marker were analysed among heterocystous cyanobacteria with maximum-parsimony, maximum-likelihood and Bayesian inference by comparing their evolutionary information to that of the 16S rRNA gene. Trees inferred from the 16S rRNA gene and the distribution of two repeat classes in the P6b stem-loop of the trnL intron were in clear conflict. The results show that, while similar heptanucleotide repeat classes I and II in the P6b stem-loop of the trnL intron could be found among distant relatives, some close relatives harboured different repeat classes with a high sequence difference. Moreover, heptanucleotide repeat class II and other sequences from the P6b stem-loop of the trnL intron interrupted several other intergenic regions in the genomes of heterocystous cyanobacteria. Cluster analyses based on conserved intron sequences without loops P6b, P9 and parts of P5 corresponded in most clades to the 16S rRNA gene phylogeny, although the relationships were not resolved well, according to low bootstrap support. Thus, the hypervariable loop sequences of the trnL intron, especially the P6b stem-loop, cannot be used for phylogenetic analysis and conclusions cannot be drawn about species relationships on the basis of these elements. Evolutionary scenarios are discussed considering the origin of the repeats.

Base Sequence↗

Bayesian analysis of level-spacing distributions for chaotic systems with broken symmetry.

Bayesian inference is applied to the nearest-neighbor and next-nearest-neighbor spacing distributions of levels of coupled superconducting microwave billiards. The weakly coupled resonators are equivalent to a quantum system with a partially broken symmetry. The coupling parameters are obtained with help from Bayes's theorem. This procedure does not require the introduction of a set of bins. The results are more accurate than those obtained from other bin-independent procedures.

Journal Article↗

Statistical mechanics of the Bayesian image restoration under spatially correlated noise.

We investigated the use of the Bayesian inference to restore noise-degraded images under conditions of spatially correlated noise. The generative statistical models used for the original image and the noise were assumed to obey multidimensional Gaussian distributions, whose covariance matrices are translational invariant. We derived an exact description to be used as the expectation for the restored image by the Fourier transformation and restored an image distorted by spatially correlated noise by using a spatially uncorrelated noise model. We found that the resulting hyperparameter estimations for the minimum error and maximal posterior marginal criteria did not coincide when the generative probabilistic model and the model used for restoration were in different classes, while they did coincide when they were in the same class.

Journal Article↗

Maximum entropy and Bayesian data analysis: Entropic prior distributions.

The problem of assigning probability distributions which reflect the prior information available about experiments is one of the major stumbling blocks in the use of Bayesian methods of data analysis. In this paper the method of maximum (relative) entropy (ME) is used to translate the information contained in the known form of the likelihood into a prior distribution for Bayesian inference. The argument is inspired and guided by intuition gained from the successful use of ME methods in statistical mechanics. For experiments that cannot be repeated the resulting "entropic prior" is formally identical with the Einstein fluctuation formula. For repeatable experiments, however, the expected value of the entropy of the likelihood turns out to be relevant information that must be included in the analysis. The important case of a Gaussian likelihood is treated in detail.

Journal Article↗