Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Phylogeny and genetic diversity of palolo worms (Palola, Eunicidae) from the tropical North Pacific and the Caribbean.

Palolo worms (Palola, Eunicidae) are best known for their annual mass spawnings, or "risings," in the South Pacific. Palola currently contains 14 morphologically similar species, mostly from shallow tropical waters. In this study, 60 specimens of Palola from nine locations in the tropical North Pacific and the Caribbean were sequenced for the two mitochondrial markers cytochrome c oxidase subunit I and 16S ribosomal RNA to infer phylogenetic relationships, genetic diversity, and phylogeography within the taxon. Phylogenetic analysis was performed using Bayesian statistics and parsimony. Vouchers of the same specimens were examined morphologically. Two major clades (A and B) can be distinguished within the monophyletic Palola. A number of individuals in clade B bear rows of ventral eyespots in the posterior body region, typical for swarming P. viridis and probably a synapomorphy for clade B. No morphological synapomorphy was found for clade A. Haplotypes from divergent clades often co-occur in the same location. Some haplotypes are geographically widespread, in one case covering the entire east-west expansion of the tropical Pacific. These results imply that despite the apparent absence of teleplanic larvae in eunicid polychaetes, long-distance dispersal is possible in at least some lineages of Palola. With the first taste of palolo I understood the Samoans' love for it. Certainly it suggested a salty caviar, but with something added, a strong, rich whiff of the mystery and fecundity of the ocean depths. -R. Steinberg. Pacific and Southeast Asian cooking. Time-Life Books, New York, 1970.

Animals↗

Comparing Bayesian estimates of genetic differentiation of molecular markers and quantitative traits: an application to Pinus sylvestris.

Comparison of the level of differentiation at neutral molecular markers (estimated as F(ST) or G(ST)) with the level of differentiation at quantitative traits (estimated as Q(ST)) has become a standard tool for inferring that there is differential selection between populations. We estimated Q(ST) of timing of bud set from a latitudinal cline of Pinus sylvestris with a Bayesian hierarchical variance component method utilizing the information on the pre-estimated population structure from neutral molecular markers. Unfortunately, the between-family variances differed substantially between populations that resulted in a bimodal posterior of Q(ST) that could not be compared in any sensible way with the unimodal posterior of the microsatellite F(ST). In order to avoid publishing studies with flawed Q(ST) estimates, we recommend that future studies should present heritability estimates for each trait and population. Moreover, to detect variance heterogeneity in frequentist methods (ANOVA and REML), it is of essential importance to check also that the residuals are normally distributed and do not follow any systematically deviating trends.

Bayes Theorem↗

IQ-NET: fast and accurate quartet phylogenetic inference using deep learning trained on empirical DNA alignments.

Phylogenetic inference is fundamental to modern biology, with many applications including evolutionary biology, epidemiology, and comparative genomics. While maximum likelihood and Bayesian methods remain the gold standard for phylogenetic analysis, they rely on simplifying assumptions and are computationally intensive. Recent machine learning approaches for phylogenetics offer speed advantages, but have several limitations: exclusive reliance on simulated data for training, inadequate handling of gaps, and sensitivity to input sequence order. Here, we introduce IQ-NET (Intelligent Quartet NETwork), a deep learning framework that solves these limitations to infer four-taxon trees. IQ-NET estimates both tree topology and branch lengths directly from gapped alignments. IQ-NET outperforms existing machine learning methods in terms of accuracy, and obtained a 24-fold speedup compared with the widely used maximum likelihood software, IQ-TREE. We finally introduce a pipeline using IQ-NET and the ASTRAL software to reconstruct a larger species tree, i.e., with more than four taxa.

Empirical data training↗

Genetic analysis of complex demographic scenarios: spatially expanding populations of the cane toad, Bufo marinus.

Inferring the spatial expansion dynamics of invading species from molecular data is notoriously difficult due to the complexity of the processes involved. For these demographic scenarios, genetic data obtained from highly variable markers may be profitably combined with specific sampling schemes and information from other sources using a Bayesian approach. The geographic range of the introduced toad Bufo marinus is still expanding in eastern and northern Australia, in each case from isolates established around 1960. A large amount of demographic and historical information is available on both expansion areas. In each area, samples were collected along a transect representing populations of different ages and genotyped at 10 microsatellite loci. Five demographic models of expansion, differing in the dispersal pattern for migrants and founders and in the number of founders, were considered. Because the demographic history is complex, we used an approximate Bayesian method, based on a rejection-regression algorithm, to formally test the relative likelihoods of the five models of expansion and to infer demographic parameters. A stepwise migration-foundation model with founder events was statistically better supported than other four models in both expansion areas. Posterior distributions supported different dynamics of expansion in the studied areas. Populations in the eastern expansion area have a lower stable effective population size and have been founded by a smaller number of individuals than those in the northern expansion area. Once demographically stabilized, populations exchange a substantial number of effective migrants per generation in both expansion areas, and such exchanges are larger in northern than in eastern Australia. The effective number of migrants appears to be considerably lower than that of founders in both expansion areas. We found our inferences to be relatively robust to various assumptions on marker, demographic, and historical features. The method presented here is the only robust, model-based method available so far, which allows inferring complex population dynamics over a short time scale. It also provides the basis for investigating the interplay between population dynamics, drift, and selection in invasive species.

Animals↗

A Bayesian approach to retransformation bias in transformed regression.

Ecological data analysis often involves fitting linear or nonlinear equations to data after transforming either the response variable, the right side of the equation, or both, so that the standard suite of regression assumptions are more closely met. However, inference is usually done in the natural metric and it is well known that retransforming back to the original metric provides a biased estimator for the mean of the response variable. For the normal linear model, fit under a log-transformation, correction factors are available to reduce this bias, but these factors may not be generally applicable to all model forms or other transformations. We demonstrate that this problem is handled in a straightforward manner using a Bayesian approach, which is general for linear and nonlinear models and other transformations and model error structures. The Bayesian framework provides a predictive distribution for the response variable so that inference can be made at the mean, or over the entire distribution to incorporate the predictive uncertainty.

Bayes Theorem↗

Monitoring of a pilot toxicity study with two adverse outcomes.

We consider monitoring a pilot toxicity study in which the adverse outcome is bivariate and the goal is to terminate the trial if evidence of excessive toxicity is encountered. We develop a Bayesian monitoring rule, based on the posterior probability that the frequency of either adverse outcome exceeds that observed under standard therapy. This rule is intuitive and ethical, and extends in a straightforward fashion from the univariate to the multivariate case. Since p-values and confidence intervals are standard methods for reporting the results of clinical trials, we also suggest how frequentist inferences may be drawn at the conclusion of a study monitored in this fashion. This work thus represents an integration of Bayesian and frequentist methodology for sequential clinical trials.

Analysis of Variance↗

Inference in linkage analysis of multifactorial traits using recombinant inbred strains of mice.

Recombinant inbred strains have been shown to be important tools for segregation and linkage analysis of multifactorial traits. Tests of association have been used as robust methods of linkage detection, however, guidelines for forming inferences from significance levels have not been generally available. In this paper, lessons learned from a Bayesian statistical approach to linkage analysis of Mendelian traits have been applied to studies of multifactorial traits. Criteria for detection of linkage based on Bonferroni's correction for multiple testing are also discussed.

Animals↗

Molecular phylogeny of the stromateoid fishes (Teleostei: Perciformes) inferred from mitochondrial DNA sequences and compared with morphology-based hypotheses.

The phylogenetic relationships among 21 species of stromateoid fishes, representing five families and 13 genera, were reconstructed using 3263bp of mitochondrial DNA sequences, including the posterior half of the 16S rRNA and entire COI and Cytb genes. The resultant molecular phylogenies were compared with previous phylogenetic hypotheses inferred from morphological characters. Molecular phylogenetic trees were constructed using the maximum parsimony, maximum likelihood, and Bayesian methods. All three methods resulted in well-resolved trees with most nodes being supported by moderate to high support values. In contrast to previous morphological analyses, which resulted in non-monophyly of Centrolophidae, all three methods utilized for the present molecular analyses supported the monophyly of Centrolophidae, as well as the reciprocal monophyly of the other stromateoid families, previous morphological hypotheses being rejected by the Templeton and Shimodaira-Hasegawa tests. In addition, the three methods indicated a sister-group relationship between Ariommatidae and Nomeidae. The position of Tetragonuridae was, however, incongruent between the MP method and the ML and Bayesian methods, being placed in the most basal position of Stromateoidei in the former, but occupying a sister relationship to Stromateidae in the latter. Comparison of the molecular phylogenies to previous morphological hypotheses suggested that evolutionary changes in morphological characters have not occurred equally among the stromateoid lineages, the evolution of the centrolophids not having been accompanied by appreciable morphological changes, whereas other stromateoids have undergone considerable morphological changes during their evolutionary history. The molecular phylogenies also shed some light on the evolutionary pattern of the pharyngeal sac, two of the four types of sac corresponding to two main lineages of Stromateoidei. Some taxonomic implications were also discussed.

Animals↗

Bottom-up and top-down dynamics in visual cortex.

A key emergent property of the primary visual cortex (V1) is the orientation selectivity of its neurons. Recent experiments demonstrate remarkable bottom-up and top-down plasticity in orientation networks of the adult cortex. The basis for such dynamics is the mechanism by which orientation tuning is created and maintained, by integration of thalamocortical and intracortical inputs. Intracellular measurements of excitatory and inhibitory synaptic conductances reveal that excitation and inhibition balance each other at all locations in the cortex. This balance is particularly critical at pinwheel centers of the orientation map, where neurons receive intracortical input from a wide diversity of local orientations. The orientation tuning of neurons in adult V1 changes systematically after short-term exposure to one stimulus orientation. Such reversible physiological shifts in tuning parallel the orientation tilt aftereffect observed psychophysically. Neurons at or near pinwheel centers show pronounced changes in orientation preference after adaptation with an oriented stimulus, while neurons in iso-orientation domains show minimal changes. Neurons in V1 of alert, behaving monkeys also exhibit short-term orientation plasticity after very brief adaptation with an oriented stimulus, on the time scale of visual fixation. Adaptation with stimuli that are orthogonal to a neuron's preferred orientation does not alter the preferred orientation but sharpens orientation tuning. Thus, successive fixation on dissimilar image patches, as happens during natural vision, combined with mechanisms of rapid cortical plasticity, actually improves orientation discrimination. Finally, natural vision involves judgements about where to look next, based on an internal model of the visual world. Experiments in behaving monkeys in which information about future stimulus locations can be acquired in one set of trials but not in another demonstrate that V1 neurons signal the acquisition of internal representations. Such Bayesian updating of responses based on statistical learning is fundamental for higher level vision, for deriving inferences about the structure of the visual world, and for the regulation of eye movements.

Adaptation, Psychological↗

Historical and contemporary multilocus population structure of Ascochyta rabiei (teleomorph: Didymella rabiei) in the Pacific Northwest of the United States.

The historical and contemporary population genetic structure of the chickpea Ascochyta blight pathogen, Ascochyta rabiei (teleomorph: Didymella rabiei), was determined in the US Pacific Northwest (PNW) using 17 putative AFLP loci, four genetically characterized, sequence-tagged microsatellite loci (STMS) and the mating type locus (MAT). A single multilocus genotype of A. rabiei (MAT1-1) was detected in 1983, which represented the first recorded appearance of Ascochyta blight of chickpea in the PNW. During the following year many additional alleles, including the other mating type allele (MAT1-2), were detected. By 1987, all alleles currently found in the PNW had been introduced. Highly significant genetic differentiation was detected among contemporary subpopulations from different hosts and geographical locations indicating restricted gene flow and/or genetic drift occurring within and among subpopulations and possible selection by host cultivar. Two distinct populations were inferred with high posterior probability which correlated to host of origin and date of sample using Bayesian model-based population structure analyses of multilocus genotypes. Allele frequencies, genotype distributions and population assignment probabilities were significantly different between the historical and contemporary samples of isolates and between isolates sampled from a resistance screening nursery and those sampled from commercial chickpea fields. A random mating model could not be rejected in any subpopulation, indicating the importance of the sexual stage of the fungus both as a source of primary inoculum for Ascochyta blight epidemics and potentially adaptive genotypic diversity.

Ascomycota↗

How meaningful are Bayesian support values?

In this study, we used an empirical example based on 100 mitochondrial genomes from higher teleost fishes to compare the accuracy of parsimony-based jackknife values with Bayesian support values. Phylogenetic analyses of 366 partitions, using differential taxon and character sampling from the entire data matrix of 100 taxa and 7,990 characters, were performed for both phylogenetic methods. The tree topology and branch-support values from each partition were compared with the tree inferred from all taxa and characters. Using this approach, we quantified the accuracy of the branch-support values assigned by the jackknife and Bayesian methods, with respect to each of 15 basal clades. In comparing the jackknife and Bayesian methods, we found that (1) both measures of support differ significantly from an ideal support index; (2) the jackknife underestimated support values; (3) the Bayesian method consistently overestimated support; (4) the magnitude by which Bayesian values overestimate support exceeds the magnitude by which the jackknife underestimates support; and (5) both methods performed poorly when taxon sampling was increased and character sampling was not increases. These results indicate that (1) the higher Bayesian support values are inappropriate (in magnitude), and (2) Bayesian support values should not be interpreted as probabilities that clades are correctly resolved. We advocate the continued use of the relatively conservative bootstrap and jackknife approaches to estimating branch support rather than the more extreme overestimates provided by the Markov Chain Monte Carlo-based Bayesian methods.

Animals↗

Reconstitution of ancestral green visual pigments of zebrafish and molecular mechanism of their spectral differentiation.

We previously reported that zebrafish have four tandemly duplicated green (RH2) opsin genes (RH2-1, RH2-2, RH2-3, and RH2-4). Absorption spectra vary widely among the four photopigments reconstituted with 11-cis retinal, with their peak absorption spectra (lambda(max)) being 467, 476, 488, and 505 nm, respectively. In this study, we inferred the ancestral amino acid (aa) sequences of the zebrafish RH2 opsins by likelihood-based Bayesian statistics and reconstituted the ancestral opsins by site-directed mutagenesis. The ancestral pigment (A1) to the four zebrafish RH2 pigments and that (A3) to RH2-3 and RH2-4 showed lambda(max) at 506 nm, while that (A2) to RH2-1 and RH2-2 showed a lambda(max) at 474 nm, indicating that a spectral shift had occurred toward the shorter wavelength on the evolutionary lineages A1 to A2 by 32 nm, A2 to RH2-1 by 7 nm, and A3 to RH2-3 by 18 nm. Pigment chimeras and site-directed mutagenesis revealed a large contribution (approximately 15 nm) of glutamic acid to glutamine substitution at residue 122 (E122Q) to the A1 to A2 and A3 to RH2-3 spectral shifts. However, the remaining spectral differences appeared to result from complex interactive effects of a number of aa replacements, each of which has only a minor spectral contribution (1-3 nm). The four zebrafish RH2 pigments cover nearly an entire range of lambda(max) distribution among vertebrate RH2 pigments and provide an excellent model to study spectral tuning mechanisms of RH2 in vertebrates.

Amino Acid Sequence↗

Single-trial lambda wave identification using a fuzzy inference system and predictive statistical diagnosis.

The aim of the study was to automate the identification of a saccade-related visual evoked potential (EP) called the lambda wave. The lambda waves were extracted from single trials of electroencephalogram (EEG) waveforms using independent component analysis (ICA). A trial was a set of EEG waveforms recorded from 64 scalp electrode locations while a saccade was performed. Forty saccade-related EEG trials (recorded from four normal subjects) were used in the study. The number of waveforms per trial was reduced from 64 to 22 by pre-processing. The application of ICA to the resulting waveforms produced 880 components (i.e. 4 subjects x 10 trials per subject x 22 components per trial). The components were divided into 373 lambda and 507 nonlambda waves by visual inspection and then they were represented by one spatial and two temporal features. The classification performance of a Bayesian approach called predictive statistical diagnosis (PSD) was compared with that of a fuzzy logic approach called a fuzzy inference system (FIS). The outputs from the two classification approaches were then combined and the resulting discrimination accuracy was evaluated. For each approach, half the data from the lambda and nonlambda wave categories were used to determine the operating parameters of the classification schemes while the rest (i.e. the validation set) were used to evaluate their classification accuracies. The sensitivity and specificity values when the classification approaches were applied to the lambda wave validation data set were as follows: for the PSD 92.51% and 91.73% respectively, for the FIS 95.72% and 89.76% respectively, and for the combined FIS and PSD approach 97.33% and 97.24% respectively (classification threshold was 0.5). The devised signal processing techniques together with the classification approaches provided for an effective extraction and classification of the single-trial lambda waves. However, as only four subjects were included, it will be valuable to further evaluate the methods on a larger group of subjects.

Adult↗

Ancestral sequence reconstruction in primate mitochondrial DNA: compositional bias and effect on functional inference.

Reconstruction of ancestral DNA and amino acid sequences is an important means of inferring information about past evolutionary events. Such reconstructions suggest changes in molecular function and evolutionary processes over the course of evolution and are used to infer adaptation and convergence. Maximum likelihood (ML) is generally thought to provide relatively accurate reconstructed sequences compared to parsimony, but both methods lead to the inference of multiple directional changes in nucleotide frequencies in primate mitochondrial DNA (mtDNA). To better understand this surprising result, as well as to better understand how parsimony and ML differ, we constructed a series of computationally simple "conditional pathway" methods that differed in the number of substitutions allowed per site along each branch, and we also evaluated the entire Bayesian posterior frequency distribution of reconstructed ancestral states. We analyzed primate mitochondrial cytochrome b (Cyt-b) and cytochrome oxidase subunit I (COI) genes and found that ML reconstructs ancestral frequencies that are often more different from tip sequences than are parsimony reconstructions. In contrast, frequency reconstructions based on the posterior ensemble more closely resemble extant nucleotide frequencies. Simulations indicate that these differences in ancestral sequence inference are probably due to deterministic bias caused by high uncertainty in the optimization-based ancestral reconstruction methods (parsimony, ML, Bayesian maximum a posteriori). In contrast, ancestral nucleotide frequencies based on an average of the Bayesian set of credible ancestral sequences are much less biased. The methods involving simpler conditional pathway calculations have slightly reduced likelihood values compared to full likelihood calculations, but they can provide fairly unbiased nucleotide reconstructions and may be useful in more complex phylogenetic analyses than considered here due to their speed and flexibility. To determine whether biased reconstructions using optimization methods might affect inferences of functional properties, ancestral primate mitochondrial tRNA sequences were inferred and helix-forming propensities for conserved pairs were evaluated in silico. For ambiguously reconstructed nucleotides at sites with high base composition variability, ancestral tRNA sequences from Bayesian analyses were more compatible with canonical base pairing than were those inferred by other methods. Thus, nucleotide bias in reconstructed sequences apparently can lead to serious bias and inaccuracies in functional predictions.

Animals↗

Uncertainty Modeling Outperforms Machine Learning for Microbiome Data Analysis.

Microbiome sequencing measures relative rather than absolute abundances, providing no direct information about total microbial load. Normalization methods attempt to compensate, but rely on strong, often untestable assumptions that can bias inference. Experimental measurements of load (e.g., qPCR, flow cytometry) offer a solution, but remain costly and uncommon. A recent high-profile study proposed that machine learning could bypass this limitation by predicting microbial load from sequencing data alone. To evaluate this claim, we assembled mutt, the largest public database of paired sequencing and load measurements, spanning 35 studies and over 15,000 samples. Using mutt, we show that published machine learning models fail to generalize: on average they perform worse than a naive baseline that always predicted the training set mean. These failures stem from covariate shift-limited shared taxa between studies, differences in community composition, and differences in preprocessing pipelines-that silently derail model inputs. In contrast, Bayesian partially identified models do not attempt to impute microbial load, but instead propagate scale uncertainty through downstream analyses. Across 30 benchmark datasets, Bayesian partially identified models consistently outperformed normalization and machine learning approaches, providing a principled and reproducible foundation for microbiome inference.

16S rRNA-seq↗

Statistical limitations in functional neuroimaging. I. Non-inferential methods and statistical models.

Functional neuroimaging (FNI) provides experimental access to the intact living brain making it possible to study higher cognitive functions in humans. In this review and in a companion paper in this issue, we discuss some common methods used to analyse FNI data. The emphasis in both papers is on assumptions and limitations of the methods reviewed. There are several methods available to analyse FNI data indicating that none is optimal for all purposes. In order to make optimal use of the methods available it is important to know the limits of applicability. For the interpretation of FNI results it is also important to take into account the assumptions, approximations and inherent limitations of the methods used. This paper gives a brief overview over some non-inferential descriptive methods and common statistical models used in FNI. Issues relating to the complex problem of model selection are discussed. In general, proper model selection is a necessary prerequisite for the validity of the subsequent statistical inference. The non-inferential section describes methods that, combined with inspection of parameter estimates and other simple measures, can aid in the process of model selection and verification of assumptions. The section on statistical models covers approaches to global normalization and some aspects of univariate, multivariate, and Bayesian models. Finally, approaches to functional connectivity and effective connectivity are discussed. In the companion paper we review issues related to signal detection and statistical inference.

Bayes Theorem↗

Bayesian statistics in medical research: an intuitive alternative to conventional data analysis.

Statistical analysis of both experimental and observational data is central to medical research. Unfortunately, the process of conventional statistical analysis is poorly understood by many medical scientists. This is due, in part, to the counter-intuitive nature of the basic tools of traditional (frequency-based) statistical inference. For example, the proper definition of a conventional 95% confidence interval is quite confusing. It is based upon the imaginary results of a series of hypothetical repetitions of the data generation process and subsequent analysis. Not surprisingly, this formal definition is often ignored and a 95% confidence interval is widely taken to represent a range of values that is associated with a 95% probability of containing the true value of the parameter being estimated. Working within the traditional framework of frequency-based statistics, this interpretation is fundamentally incorrect. It is perfectly valid, however, if one works within the framework of Bayesian statistics and assumes a 'prior distribution' that is uniform on the scale of the main outcome variable. This reflects a limited equivalence between conventional and Bayesian statistics that can be used to facilitate a simple Bayesian interpretation based on the results of a standard analysis. Such inferences provide direct and understandable answers to many important types of question in medical research. For example, they can be used to assist decision making based upon studies with unavoidably low statistical power, where non-significant results are all too often, and wrongly, interpreted as implying 'no effect'. They can also be used to overcome the confusion that can result when statistically significant effects are too small to be clinically relevant. This paper describes the theoretical basis of the Bayesian-based approach and illustrates its application with a practical example that investigates the prevalence of major cardiac defects in a cohort of children born using the assisted reproduction technique known as ICSI (intracytoplasmic sperm injection).

Bayes Theorem↗

Evolutionary rates, divergence dates, and the performance of mitochondrial genes in Bayesian phylogenetic analysis.

The mitochondrial genome is one of the most frequently used loci in phylogenetic and phylogeographic analyses, and it is becoming increasingly possible to sequence and analyze this genome in its entirety from diverse taxa. However, sequencing the entire genome is not always desirable or feasible. Which genes should be selected to best infer the evolutionary history of the mitochondria within a group of organisms, and what properties of a gene determine its phylogenetic performance? The current study addresses these questions in a Bayesian phylogenetic framework with reference to a phylogeny of plethodontid and related salamanders derived from 27 complete mitochondrial genomes; this topology is corroborated by nuclear DNA and morphological data. Evolutionary rates for each mitochondrial gene and divergence dates for all nodes in the plethodontid mitochondrial genome phylogeny were estimated in both Bayesian and maximum likelihood frameworks using multiple fossil calibrations, multiple data partitions, and a clock-independent approach. Bayesian analyses of individual genes were performed, and the resulting trees compared against the reference topology. Ordinal logistic regression analysis of molecular evolution rate, gene length, and the G-shape parameter a demonstrated that slower rate of evolution and longer gene length both increased the probability that a gene would perform well phylogenetically. Estimated rates of molecular evolution vary 84-fold among different mitochondrial genes and different salamander lineages, and mean rates among genes vary 15-fold. Despite having conserved amino acid sequences, cox1, cox2, cox3, and cob have the fastest mean rates of nucleotide substitution, and the greatest variation in rates, whereas rrnS and rrnL have the slowest rates. Reasons underlying this rate variation are discussed, as is the extensive rate variation in cox1 in light of its proposed role in DNA barcoding.

Animals↗