Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Inference from Inadequate and Inaccurate Data, III.

Having measured D numerical properties of a physical object E which requires many more than D parameters for its complete description, we want to estimate P other numerical properties of E. Continuing the discussion in papers I(1) and II,(2) the present paper gives estimates when we believe it likely that we can guess an upper bound M on the Hilbert norm not of h(E), the model representing E in some Hilbert space, but of the orthogonal projection of h(E) onto a sufficiently large subspace. In addition, the present paper simplifies the notation of I, and makes explicit the application of Bayesian subjective probability when there are errors in the data and we want to find the joint probability distribution of more than one prediction.

Journal Article↗

Bayesian analysis of multilocus association in quantitative and qualitative traits.

A Bayesian model-based method for multilocus association analysis of quantitative and qualitative (binary) traits is presented. The method selects a trait-associated subset of markers among candidates, and is equally applicable for analyzing wide chromosomal segments (genome scans) and small candidate regions. The method can be applied in situations involving missing genotype data. The number of trait loci, their marker positions, and the magnitudes of their gene effects (strengths of association) are all estimated simultaneously. The inference of parameters is based on their posterior distributions, which are obtained through Markov chain Monte Carlo simulations. The strengths of the approach are: 1) flexible use of oligogenic models with unknown number of loci, 2) performing the estimation of association jointly with model selection, and 3) avoidance of the multiple testing problem, which typically complicates the approaches based on association testing. The performance of the method was tested and compared to the multilocus conditional search procedure by analyzing two simulated data sets. We also applied the method to cystic fibrosis haplotype data (two-locus haplotypes), where gene position has already been identified. The method is implemented as a software package, which is freely available for research purposes under the name BAMA.

Algorithms↗

Measuring gametic disequilibrium from multilocus data.

We describe a Bayesian approach to analyzing multilocus genotype or haplotype data to assess departures from gametic (linkage) equilibrium. Our approach employs a Markov chain Monte Carlo (MCMC) algorithm to approximate the posterior probability distributions of disequilibrium parameters. The distributions are computed exactly in some simple settings. Among other advantages, posterior distributions can be presented visually, which allows the uncertainties in parameter estimates to be readily assessed. In addition, background knowledge can be incorporated, where available, to improve the precision of inferences. The method is illustrated by application to previously published datasets; implications for multilocus forensic match probabilities and for simple association-based gene mapping are also discussed.

Algorithms↗

Mitochondrial and nuclear DNA phylogeography of European grayling (Thymallus thymallus): evidence for secondary contact zones in central Europe.

Mitochondrial and microsatellite DNA markers were applied to infer the phylogeography, intraspecific diversity and dynamics of the distributional history of European grayling (Thymallus thymallus) with focus on its central and northern European distribution range. Phylogenetic and nested clade analyses revealed at least four major mtDNA lineages, which evolved in geographical isolation during the Pleistocene. These lineages should be recognized as the basic evolutionary significant units (ESUs) for grayling in central and northern Europe. In addition, and in contrast to previous work on grayling, the results of Bayesian analysis of individual admixture coefficients, two-dimensional scaling analysis and spatial analysis of molecular variance provided evidence for a high level of admixture among major lineages in contact zones between drainages (e.g. the low mountain range of Germany), most likely resulting from glacial perturbations and ancient river connections between drainages during the Pleistocene glaciations. Even within river systems, a high level of differentiation among populations was revealed as indicated by the microsatellite data. Grayling sampled from 29 sites displayed high levels of differentiation (overall F(ST) = 0.367), a high number of private alleles and high bootstrap support for the genetic distance-based population clusters across 12 loci. We specifically discuss our results in context of phylogeograpic studies on other European freshwater fish species with habitat preferences similar to those of grayling. Our study shows that both large-scale phylogeographical and detailed genetic analyses on a fine scale are mandatory for developing appropriate conservation guidelines of endangered species.

Animals↗

Molecular phylogeny of acantharian and polycystine radiolarians based on ribosomal DNA sequences, and some comparisons with data from the fossil record.

Polycystines (spumellarians, nassellarians, and collodarians), phaeodarians, and acantharians are marine planktonic protists that have been conventionally and collectively called "radiolaria". Recent molecular phylogenetic studies revealed radiolarian polyphyly with phaeodarians being a separate offshoot. Collodarians and nassellarians are also shown to form a monophyletic group, but other aspects of radiolarian phylogeny, such as interrelations among polycystines and acantharians, remained uncertain. Here, we present molecular phylogenetic analyses including new ribosomal RNA sequences from ten spumellarians and nine nassellarians, based on Bayesian and maximum-likelihood methods. Results indicate that the Polycystinea is a paraphyletic group, with Bayesian analysis suggesting that spumellarians form a clade with acantharians. The heliozoan-like protist Sticholonche appears as a sister to the spumellarian clade. The nassellarian Eucyrtidium is located outside the clade including the other nassellarians and collodarians. The mineralogy of the test of extant radiolarians and the tree topology obtained in this work suggest that acantharians and spumellarians evolved from an ancestor with a siliceous skeleton. Collodarians and nassellarians form a well-supported clade and one might infer from the fossil record that they may have diverged between the Jurassic and the Eocene.

Animals↗

Using approximate Bayesian computation to estimate tuberculosis transmission parameters from genotype data.

Tuberculosis can be studied at the population level by genotyping strains of Mycobacterium tuberculosis isolated from patients. We use an approximate Bayesian computational method in combination with a stochastic model of tuberculosis transmission and mutation of a molecular marker to estimate the net transmission rate, the doubling time, and the reproductive value of the pathogen. This method is applied to a published data set from San Francisco of tuberculosis genotypes based on the marker IS6110. The mutation rate of this marker has previously been studied, and we use those estimates to form a prior distribution of mutation rates in the inference procedure. The posterior point estimates of the key parameters of interest for these data are as follows: net transmission rate, 0.69/year [95% credibility interval (C.I.) 0.38, 1.08]; doubling time, 1.08 years (95% C.I. 0.64, 1.82); and reproductive value 3.4 (95% C.I. 1.4, 79.7). These figures suggest a rapidly spreading epidemic, consistent with observations of the resurgence of tuberculosis in the United States in the 1980s and 1990s.

Algorithms↗

Phylogeny of the photosynthetic euglenophytes inferred from the nuclear SSU and partial LSU rDNA.

Previous studies using the nuclear SSU rDNA have indicated that the photosynthetic euglenoids are a monophyletic group; however, some of the genera within the photosynthetic lineage are not monophyletic. To test these results further, evolutionary relationships among the photosynthetic genera were investigated by obtaining partial LSU nuclear rDNA sequences. Taxa from each of the external clades of the SSU rDNA-based phylogeny were chosen to create a combined dataset and to compare the individual LSU and SSU rDNA datasets. Conserved areas of the aligned sequences for both the LSU and SSU rDNA were used to generate parsimony, log-det, maximum-likelihood and Bayesian trees. The SSU and LSU rDNA consistently generated the same seven terminal clades; however, the relationship among those clades varied depending on the type of analysis and the dataset used. The combined dataset generated a more robust phylogeny, but the relationships among clades still varied. The addition of the LSU rDNA dataset to the euglenophyte phylogeny supports the view that the genera Euglena, Lepocinclis and Phacus are not monophyletic and substantiates the existence of several well-supported clades. A secondary structural model for the D2 region of the LSU rDNA was proposed on the basis of compensatory base changes found in the alignment.

Animals↗

Assays for recombinant proteins: a problem in non-linear calibration.

Quantification of protein levels in biological matrices such as serum or plasma frequently relies on the techniques of immunoassay or bioassay. The relevant statistical problem is that of non-linear calibration, where one estimates analyte concentration in an unknown sample from a calibration curve fit to known standard concentrations. This paper discusses a general framework for calibration curve fit to known standard concentrations. This paper discusses a general framework for calibration inference, that of the non-linear mixed effects model. Within this framework, we consider two issues in depth: accurate characterization of intra-assay variation, and the use of empirical Bayes methods in calibration. We show that proper characterization of intra-assay variability requires pooling of information across several assay runs. Simulation work indicates that use of empirical Bayes methods may afford considerable gain in efficiency; one must weigh this gain against practical considerations in the implementation of Bayesian techniques. We illustrate the methods discussed using a cell-based bioassay for the recombinant hormone relaxin.

Bayes Theorem↗

Bayesian mapping of quantitative trait loci under complicated mating designs.

Quantitative trait loci (QTL) are easily studied in a biallelic system. Such a system requires the cross of two inbred lines presumably fixed for alternative alleles of the QTL. However, development of inbred lines can be time consuming and cost ineffective for species with long generation intervals and severe inbreeding depression. In addition, restriction of the investigation to a biallelic system can sometimes be misleading because many potentially important allelic interactions do not have a chance to express and thus fail to be detected. A complicated mating design involving multiple alleles mimics the actual breeding system. However, it is difficult to develop the statistical model and algorithm using the classical maximum-likelihood method. In this study, we investigate the application of a Bayesian method implemented via the Markov chain Monte Carlo (MCMC) algorithm to QTL mapping under arbitrarily complicated mating designs. We develop the method under a mixed-model framework where the genetic values of founder alleles are treated as random and the nongenetic effects are treated as fixed. With the MCMC algorithm, we first draw the gene flows from the founders to the descendants for each QTL and then draw samples of the genetic parameters. Finally, we are able to simultaneously infer the posterior distribution of the number, the additive and dominance variances, and the chromosomal locations of all identified QTL.

Analysis of Variance↗

Molecular phylogeny of Sinocyclocheilus (Cypriniformes: Cyprinidae) inferred from mitochondrial DNA sequences.

More than 10 species within the freshwater fish genus Sinoncyclocheilus adapt to caves and show different degrees of degeneration of eyes and pigmentation. Therefore, this genus can be useful for studying evolutionary developmental mechanisms, role of natural selection and adaptation in cave animals. To better understand these processes, it is indispensable to have background knowledge about phylogenetic relationships of surface and cave species within this genus. To investigate phylogenetic relationships among species within this genus, we determined nucleotide sequences of complete mitochondrial cytochrome b gene (1140 bp) and partial ND4 gene (1032 bp) of 31 recognized ingroup species and one outgroup species Barbodes laticeps. Phylogenetic trees were reconstructed using maximum parsimony, Bayesian, and maximum likelihood analyses. Our phylogenetic results showed that all species except for two surface species S. jii and S. macrolepis clustered as five major monophyletic clades (I, II, III, IV, and V) with strong supports. S. jii was the most basal species in all analyses, but the position of S. macrolepis was not resolved. The cave species were polyphyletic and occurred in these five major clades. Our results indicate that adaptation to cave environments has occurred multiple times during the evolutionary history of Sinocyclocheilus. The branching orders among the clades I, II, III, and IV were not resolved, and this might be due to early rapid radiation in Sinocyclocheilus. All species distributed in Yunnan except for S. rhinocerous and S. hyalinus formed a strongly supported monophyletic group (clade V), probably reflecting their common origins. This result suggested that the diversification of Sinocyclocheilus in Yunnan may correlate with the uplifting of Yunnan Plateau.

Adaptation, Physiological↗

Introgression among maternal lineages inferred from complete mitogenomes and molecular dating helps resolve phylogeography of European roe deer.

BACKGROUND: The European roe deer (Capreolus capreolus) is one of the most widespread ungulates in Europe, with a phylogeographic structure mainly shaped by Pleistocene glacial cycles and secondary contacts with the Siberian roe deer (C. pygargus). METHODS: We sequenced 52 complete mitogenomes of C. capreolus from Slovenia, Poland and France, and combined them with 24 publicly available sequences of C. capreolus and C. pygargus, yielding an alignment of 76 genomes representing 59 haplotypes (42 from C. capreolus and 17 from C. pygargus). Phylogeographic structure was assessed using a median-joining network, and divergence times were estimated using a time-calibrated Bayesian phylogeny based on mitochondrial coding regions, incorporating published ancient C. pygargus mitogenomes. We additionally screened mitochondrial protein-coding genes for selection. RESULTS: The haplotype network recovered the three major European roe deer clades (Eastern, Central, and Western) and detected Central-clade haplotypes in France. Two Polish haplotypes (Cp9 and Cp10), detected in C. capreolus, clustered within the C. pygargus mitochondrial lineage, supporting mitochondrial introgression. Time-calibrated phylogenies placed introgressed haplotypes within established C. pygargus lineages. Selection analyses provided limited evidence for episodic positive selection restricted to a small number of codons. CONCLUSIONS: Whole mitogenomes improve resolution of roe deer phylogeography and reveal introgressed maternal lineages, while time-calibrated phylogenies and selection tests add evolutionary context for interpreting mtDNA diversity in genus Capreolus.

Animals↗

Linkage analysis of quantitative trait loci in multiple line crosses.

Simple line crosses, for example, backcross and F2, are commonly used in mapping quantitative trait loci (QTL). However, these simple crosses are rarely used alone in commercial plant breeding; rather, crosses involving multiple inbred lines or several simple crosses but connected by shared inbred lines may be common in plant breeding. Mapping QTL using crosses of multiple lines is more relevant to plant breeding. Unfortunately, current statistical methods and computer programs of QTL mapping are all designed for simple line crosses or multiple line crosses but under a regular mating system. It is not straightforward to extend the existing methods to handle multiple line crosses under irregular and complicated mating designs. The major hurdle comes from irregular inbreeding, multiple generations, and multiple alleles. In this study, we develop a Bayesian method implemented via the Markov chain Monte Carlo (MCMC) algorithm for mapping QTL using complicated multiple line crosses. With the MCMC algorithm, we are able to draw a complete path of the gene flow from founder alleles to their descendents via a recursive process. This has greatly simplified the problem caused by irregular mating and inbreeding in the mapping population. Adopting the reversible jump MCMC algorithm, we are able to simultaneously search for multiple QTL along the genome. We can even infer the posterior distribution of the number of QTL, one of the most important parameters in QTL study. Application of the new MCMC based QTL mapping procedure is demonstrated using two different mating designs. Design I involves two inbred lines and their derived F1, F2, and BC populations. Design II is a half-diallel cross involving three inbred lines. The two designs appear different, but can be handled with the same robust computer program.

Algorithms↗

Projections of lung cancer mortality in West Germany: a case study in Bayesian prediction.

We apply a generalized Bayesian age-period-cohort (APC) model to a data-set on lung cancer mortality in West Germany, in the period 1952-1996. Our goal is to predict future death rates until the year 2010, separately for males and females. Since age and period are not measured on the same grid, we propose a generalized APC model where consecutive cohort parameters represent strongly overlapping birth cohorts. This approach results in a rather large number of parameters, where standard algorithms for statistical inference by Markov chain Monte Carlo methods turn out to be computationally intensive. We propose a more efficient implementation based on ideas of block sampling from the time series literature. We entertain two different formulations, penalizing either first or second differences of age, period and cohort parameters. To assess the predictive quality of both formulations, we first forecast the rates for the period 1987-1996 based on data until 1986. A comparison with the actual observed rates is made based on a predictive deviance criterion. Predictions of lung cancer mortality until 2010 are then reported and a modification of the formulation in order to include information on cigarette consumption is finally described.To whom correspondence should be addressed. Currently at Imperial College School of Medicine, Department of Epidemiology and Public Health, Norfolk Place, London W2 1PG, UK.

Journal Article↗

Genomic Tracking of Market-Derived Bull Shark Fins Back to Source Population of Origin.

International trade of shark fins remains difficult to monitor because products are rarely labelled to species and are often highly processed, resulting in severely degraded DNA. For several shark species listed under Appendix II of the Convention on International Trade in Endangered Species of Wild Fauna and Flora (CITES), this limits external verification of source populations supplying global trade hubs. Here, we assess whether nuclear genomic approaches can be applied to market-derived bull shark (Carcharhinus leucas) fins to determine their population of origin. We analysed dried fin trimmings collected from retail vendors in Hong Kong SAR, one of the world's largest dried shark fin trade hubs, using a targeted DArTcap single nucleotide polymorphism (SNP) panel, originally developed for population genomic studies of this species. Despite substantial DNA degradation, genomic libraries were successfully obtained for most samples, yielding sufficient SNP data to perform robust provenance and sex assignment. Using a Bayesian mixed-stock analysis, most fin samples were assigned to the Indo-West Pacific (71.4%), with smaller contributions from the western Atlantic (22.6%) and eastern Pacific (3.0%). Genetic sex assignment revealed twice as many males as females, although results indicated a conservative bias towards male assignment due to the limited number of X-linked markers available in degraded samples. Our results demonstrate that genome-wide targeted approaches can be effectively applied to highly processed shark fin products to infer population sources and sex composition. This study provides proof-of-concept for integrating genomics into shark trade monitoring, highlighting its potential to improve traceability, support CITES implementation and inform conservation and fisheries management, particularly for species with well-resolved population structure.

Animals↗

A nuclear phylogeny of the Florideophyceae (Rhodophyta) inferred from combined EF2, small subunit and large subunit ribosomal DNA: establishing the new red algal subclass Corallinophycidae.

Previous studies have indicated that resolution of supraordinal relationships in the red algal class Florideophyceae will require additional characters, improved taxon sampling and optimized methods of phylogenetic analysis. To this end, we have generated data to introduce a novel nuclear marker to red algal systematics, elongation factor 2, as well as expanded ribosomal DNA alignments (SSU and LSU) to include 62 ingroup and 4 outgroup taxa. Both single gene and combined data sets were considered. Our analyses resulted in better resolution of both deep as well as more recent divergences, with higher support realized at many nodes. Distance, parsimony and bayesian analyses of the single gene and combined data sets indicated that the subclasses Hildenbrandiophycidae, Ahnfeltiophycidae and Rhodymeniophycidae were monophyletic, whereas the Nemaliophycidae was polyphyletic: one lineage containing the Rhodogorgonales and Corallinales (CR complex); and the other containing the Acrochaetiales, Balbianiales, Balliales, Batrachospermales, Colaconematales, Nemaliales, Palmariales, and Thoreales (APB complex). Based on these results a new subclass of the Florideophyceae, the Corallinophycidae subclassis nov., is proposed to accommodate the Corallinales and Rhodogorgonales. In addition to resolving supraordinal relationships, the present analyses resolved some novel ordinal affinities within the Nemaliophycidae and Rhodymeniophycidae, which are discussed here.

Cell Nucleus↗

Inferring global levels of alternative splicing isoforms using a generative model of microarray data.

MOTIVATION: Alternative splicing (AS) is a frequent step in metozoan gene expression whereby the exons of genes are spliced in different combinations to generate multiple isoforms of mature mRNA. AS functions to enrich an organism's proteomic complexity and regulates gene expression. Despite its importance, the mechanisms underlying AS and its regulation are not well understood, especially in the context of global gene expression patterns. We present here an algorithm referred to as the Generative model for the Alternative Splicing Array Platform (GenASAP) that can predict the levels of AS for thousands of exon skipping events using data generated from custom microarrays. GenASAP uses Bayesian learning in an unsupervised probability model to accurately predict AS levels from the microarray data. GenASAP is capable of learning the hybridization profiles of microarray data, while modeling noise processes and missing or aberrant data. GenASAP has been successfully applied to the global discovery and analysis of AS in mammalian cells and tissues. RESULTS: GenASAP was applied to data obtained from a custom microarray designed for the monitoring of 3126 AS events in mouse cells and tissues. The microarray design included probes specific for exon body and junction sequences formed by the splicing of exons. Our results show that GenASAP provides accurate predictions for over one-third of the total events, as verified by independent RT-PCR assays. SUPPLEMENTARY INFORMATION: http://www.psi.toronto.edu/GenASAP.

Algorithms↗

Building statistical models to analyze species distributions.

Models of the geographic distributions of species have wide application in ecology. But the nonspatial, single-level, regression models that ecologists have often employed do not deal with problems of irregular sampling intensity or spatial dependence, and do not adequately quantify uncertainty. We show here how to build statistical models that can handle these features of spatial prediction and provide richer, more powerful inference about species niche relations, distributions, and the effects of human disturbance. We begin with a familiar generalized linear model and build in additional features, including spatial random effects and hierarchical levels. Since these models are fully specified statistical models, we show that it is possible to add complexity without sacrificing interpretability. This step-by-step approach, together with attached code that implements a simple, spatially explicit, regression model, is structured to facilitate self-teaching. All models are developed in a Bayesian framework. We assess the performance of the models by using them to predict the distributions of two plant species (Proteaceae) from South Africa's Cape Floristic Region. We demonstrate that making distribution models spatially explicit can be essential for accurately characterizing the environmental response of species, predicting their probability of occurrence, and assessing uncertainty in the model results. Adding hierarchical levels to the models has further advantages in allowing human transformation of the landscape to be taken into account, as well as additional features of the sampling process.

Bayes Theorem↗

A Bayesian approach to jointly estimate centre and treatment by centre heterogeneity in a proportional hazards model.

When multicentre clinical trial data are analysed, it has become more and more popular to look for possible heterogeneity in outcome between centres. However, beyond the investigation of such heterogeneity, it is also interesting to consider heterogeneity in treatment effect over centres. For time-to-event outcomes, this may be investigated by including a random centre effect and a random treatment by centre interaction in a Cox proportional hazards model. Assuming independence between the random effects, we propose a Bayesian approach to fit our proposed model. The parameters of interest are the variance components sigma(0) (2) and sigma(1) (2) of these random effects, which can be interpreted as a measure of centre and treatment effect over centres heterogeneity of the hazard. These variance components are estimated from their marginal posterior density after integrating out the fixed treatment effect and the random effects. As this integration cannot be performed analytically, the marginal posterior density is approximated using the Laplace integration technique. Statistical inference is then based on the characteristics of the posterior marginal density, such as the mode and the standard deviation. We demonstrate the proposed technique using data from a pooled database of seven EORTC bladder cancer clinical trials. Substantial centre and treatment effect over centres heterogeneity in disease-free interval was found.

Bayes Theorem↗