Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Molecular phylogeny of the benthic shallow-water octopuses (Cephalopoda: Octopodinae).

Octopus has been regarded as a "catch all" genus, yet its monophyly is questionable and has been untested. We inferred a broad-scale phylogeny of the benthic shallow-water octopuses (subfamily Octopodinae) using amino acid sequences of two mitochondrial DNA genes: Cytochrome oxidase subunit III and Cytochrome b apoenzyme, and the nuclear DNA gene Elongation Factor-1alpha. Sequence data were obtained from 26 Octopus species and from four related genera. Maximum likelihood and Bayesian approaches were implemented to estimate the phylogeny, and non-parametric bootstrapping was used to verify confidence for Bayesian topologies. Phylogenetic relationships between closely related species were generally well resolved, and groups delineated, but the genes did not resolve deep divergences well. The phylogenies indicated strongly that Octopus is not monophyletic, but several monophyletic groups were identified within the genus. It is therefore clear that octopodid systematics requires major revision.

Amino Acid Sequence↗

Appraisal of the consequences of the DDT-induced bottleneck on the level and geographic distribution of neutral genetic variation in Canadian peregrine falcons, Falco peregrinus.

Peregrine falcon populations underwent devastating declines in the mid-20th century due to the bioaccumulation of organochlorine contaminants, becoming essentially extirpated east of the Great Plains and significantly reduced elsewhere in North America. Extensive re-introduction programs and restrictions on pesticide use in Canada and the United States have returned many populations to predecline sizes. A proper population genetic appraisal of the consequences of this decline requires an appropriate context defined by (i) meaningful demographic entities; and (ii) suitable reference populations. Here we explore the validity of currently recognized subspecies designations using data from the mitochondrial control region and 11 polymorphic microsatellite loci taken from 184 contemporary individuals from across the breeding range, and compare patterns of population genetic structure with historical patterns inferred from 95 museum specimens. Of the three North American subspecies, the west coast marine subspecies Falco peregrinus pealei is well differentiated genetically in both time periods using nuclear loci. In contrast, the partitioning of continental Falco peregrinus anatum and arctic Falco peregrinus tundrius subspecies is not substantiated, as individuals from these subspecies are historically indistinguishable genetically. Bayesian clustering analyses demonstrate that contemporary genetic differentiation between these two subspecies is mainly due to changes within F. p. anatum (specifically the southern F. p. anatum populations). Despite expectations and a variety of tests, no genetic bottleneck signature is found in the identified populations; in fact, many contemporary indices of diversity are higher than historical values. These results are rationalized by the promptness of the recovery and the possible introduction of new genetic material.

Animals↗

The statistical analysis of circadian phase and amplitude in constant-routine core-temperature data.

Accurate estimation of the phases and amplitude of the endogenous circadian pacemaker from constant-routine core-temperature series is crucial for making inferences about the properties of the human biological clock from data collected under this protocol. This paper presents a set of statistical methods based on a harmonic-regression-plus-correlated-noise model for estimating the phases and the amplitude of the endogenous circadian pacemaker from constant-routine core-temperature data. The methods include a Bayesian Monte Carlo procedure for computing the uncertainty in these circadian functions. We illustrate the techniques with a detailed study of a single subject's core-temperature series and describe their relationship to other statistical methods for circadian data analysis. In our laboratory, these methods have been successfully used to analyze more than 300 constant routines and provide a highly reliable means of extracting phase and amplitude information from core-temperature data.

Adult↗

Fossil calibrations and molecular divergence time estimates in centrarchid fishes (Teleostei: Centrarchidae).

Molecular clock methods allow biologists to estimate divergence times, which in turn play an important role in comparative studies of many evolutionary processes. It is well known that molecular age estimates can be biased by heterogeneity in rates of molecular evolution, but less attention has been paid to the issue of potentially erroneous fossil calibrations. In this study we estimate the timing of diversification in Centrarchidae, an endemic major lineage of the diverse North American freshwater fish fauna, through a new approach to fossil calibration and molecular evolutionary model selection. Given a completely resolved multi-gene molecular phylogeny and a set of multiple fossil-inferred age estimates, we tested for potentially erroneous fossil calibrations using a recently developed fossil cross-validation. We also used fossil information to guide the selection of the optimal molecular evolutionary model with a new fossil jackknife method in a fossil-based model cross-validation. The centrarchid phylogeny resulted from a mixed-model Bayesian strategy that included 14 separate data partitions sampled from three mtDNA and four nuclear genes. Ten of the 31 interspecific nodes in the centrarchid phylogeny were assigned a minimal age estimate from the centrarchid fossil record. Our analyses identified four fossil dates that were inconsistent with the other fossils, and we removed them from the molecular dating analysis. Using fossil-based model cross-validation to determine the optimal smoothing value in penalized likelihood analysis, and six mutually consistent fossil calibrations, the age of the most recent common ancestor of Centrarchidae was 33.59 million years ago (mya). Penalized likelihood analyses of individual data partitions all converged on a very similar age estimate for this node, indicating that rate heterogeneity among data partitions is not confounding our analyses. These results place the origin of the centrarchid radiation at a time of major faunal turnover as the fossil record indicates that the most diverse lineages of the North American freshwater fish fauna originated at the Eocene-Oligocene boundary, approximately 34 mya. This time coincided with major global climate change from warm to cool temperatures and a signature of elevated lineage extinction and origination in the fossil record across the tree of life. Our analyses demonstrate the utility of fossil cross-validation to critically assess individual fossil calibration points, providing the ability to discriminate between consistent and inconsistent fossil age estimates that are used for calibrating molecular phylogenies.

Animals↗

Genetic diversity within the Albugo candida complex (Peronosporales, Oomycota) inferred from phylogenetic analysis of ITS rDNA and COX2 mtDNA sequences.

Albugo candida is a destructive fungus infecting brassicaceous hosts. The genetic diversity within the A. candida complex from various host plants was investigated by sequence analysis of the internal transcribed spacer (ITS) region of rDNA and the cytochrome c oxidase subunit II (COX2) region of mtDNA. The aligned nucleotide sequences of A. candida shared significantly high distances, up to 20.4 and 8.9%, in two genes. The phylogenetic trees, obtained using the Bayesian method and maximum parsimony analysis, showed two separate groups that corresponded to the host genera. Group I included A. candida isolates infecting Arabis, Autrieta, Berteroa, Biscutella, Brassica, Cardaminopsis, Diplotaxis, Eruca, Erysimum, Heliophila, Iberis, Lunaria, Raphanus, Sinapis, Sisymbrium, and Thlaspi. Group II contained all isolates from Capsella, Descurainia, Diptychocarpus, Draba, and Lepidium. The genetic similarities between the two genes among isolates within Group I were 99.0-100% and 99.6-100%, while those within Group II were 90.4-100% and 91.1-100%, respectively, showing considerably lower values than for Group I. The A. candida isolates from Capsella bursa-pastoris in Korea are clearly separated by sequence analysis for the two genes compared to those from Wales, England, and the USA. Based on the molecular data from the two genes, we suggest the high degree of genetic diversity exhibited within A. candida complexes warrants their division into several distinct species.

Base Sequence↗

Molecular phylogenetics and asexuality in the brine shrimp Artemia.

Explaining cases of long-term persistence of parthenogenesis has proven an arduous task for evolutionary biologists. Interpreting sexual-asexual interactions though has recently advanced owing to methodological design, increased taxon sampling and choice of model organisms. We inferred the phylogeny of Artemia, a halophilic branchiopod genus of sexual and parthenogenetic forms with cosmopolitan distribution, marked geographic patterns and ecological partitioning. Joint analysis of newly derived ITS1 sequences and 16S RFLP markers from global isolates indicates significant interspecific divergence as well as pronounced diversity for parthenogens, matching that of sexual ancestors. Maximum parsimony, maximum likelihood, and Bayesian methods were largely congruent in reconstructing the phylogeny of the genus. Given the current sampling, at least four independent origins of parthenogenesis are deduced. Molecular clock calibrations based on biogeographic landmarks indicate that the lineage leading to A. persimilis diverged from the common ancestor of all Artemia species between 80 and 90 MYA at the time of separation of Africa from South America, whereas parthenogenesis first appeared at least 3 MYA. Common mitochondrial DNA haplotypes delineate A. urmiana and A. tibetiana as possible maternal parents of several clonal lineages. A novel topological placement of A. franciscana as a sister clade to all Asian Artemia and parthenogenetic forms is proposed and also supported by ITS1 length and other existing data.

Animals↗

ConSurf 2005: the projection of evolutionary conservation scores of residues on protein structures.

Key amino acid positions that are important for maintaining the 3D structure of a protein and/or its function(s), e.g. catalytic activity, binding to ligand, DNA or other proteins, are often under strong evolutionary constraints. Thus, the biological importance of a residue often correlates with its level of evolutionary conservation within the protein family. ConSurf (http://consurf.tau.ac.il/) is a web-based tool that automatically calculates evolutionary conservation scores and maps them on protein structures via a user-friendly interface. Structurally and functionally important regions in the protein typically appear as patches of evolutionarily conserved residues that are spatially close to each other. We present here version 3.0 of ConSurf. This new version includes an empirical Bayesian method for scoring conservation, which is more accurate than the maximum-likelihood method that was used in the earlier release. Various additional steps in the calculation can now be controlled by a number of advanced options, thus further improving the accuracy of the calculation. Moreover, ConSurf version 3.0 also includes a measure of confidence for the inferred amino acid conservation scores.

Amino Acid Substitution↗

Phylogenetic status and matrilineal structure of the biting midge, Culicoides imicola, in Portugal, Rhodes and Israel.

The biting midge Culicoides imicola Kieffer (Diptera: Ceratopogonidae) is the most important Old World vector of African horse sickness (AHS) and bluetongue (BT). Recent increases of BT incidence in the Mediterranean basin are attributed to its increased abundance and distribution. The phylogenetic status and genetic structure of C. imicola in this region are unknown, despite the importance of these aspects for BT epidemiology in the North American BT vector. In this study, analyses of partial mitochondrial cytochrome oxidase subunit I gene (COI) sequences were used to infer phylogenetic relationships among 50 C. imicola from Portugal, Rhodes, Israel, and South Africa and four other species of the Imicola Complex from southern Africa, and to estimate levels of matrilineal subdivision in C. imicola between Portugal and Israel. Eleven haplotypes were detected in C. imicola, and these formed one well-supported clade in maximum likelihood and Bayesian trees implying that the C. imicola samples comprise one phylogenetic species. Molecular variance was distributed mainly between Portugal and Israel, with no haplotypes shared between these countries, suggesting that female-mediated gene flow at this scale has been either limited or non-existent. Our results provide phylogenetic evidence that C. imicola in the study areas are potentially competent AHS and BT vectors. The geographical structure of the C. imicola COI haplotypes was concordant with that of BT virus serotypes in recent BT outbreaks in the Mediterranean basin, suggesting that population subdivision in its vector can impose spatial constraints on BT virus transmission.

African Horse Sickness↗

The molecular population genetics of HIV-1 group O.

HIV-1 group O originated through cross-species transmission of SIV from chimpanzees to humans and has established a relatively low prevalence in Central Africa. Here, we infer the population genetics and epidemic history of HIV-1 group O from viral gene sequence data and evaluate the effect of variable evolutionary rates and recombination on our estimates. First, model selection tools were used to specify suitable evolutionary and coalescent models for HIV group O. Second, divergence times and population genetic parameters were estimated in a Bayesian framework using Markov chain Monte Carlo sampling, under both strict and relaxed molecular clock methods. Our results date the origin of the group O radiation to around 1920 (1890-1940), a time frame similar to that estimated for HIV-1 group M. However, group O infections, which remain almost wholly restricted to Cameroon, show a slower rate of exponential growth during the twentieth century, explaining their lower current prevalence. To explore the effect of recombination, the Bayesian framework is extended to incorporate multiple unlinked loci. Although recombination can bias estimates of the time to the most recent common ancestor, this effect does not appear to be important for HIV-1 group O. In addition, we show that evolutionary rate estimates for different HIV genes accurately reflect differential selective constraints along the HIV genome.

Bayes Theorem↗

Identifying protein complexes in high-throughput protein interaction screens using an infinite latent feature model.

We propose a Bayesian approach to identify protein complexes and their constituents from high-throughput protein-protein interaction screens. An infinite latent feature model that allows for multi-complex membership by individual proteins is coupled with a graph diffusion kernel that evaluates the likelihood of two proteins belonging to the same complex. Gibbs sampling is then used to infer a catalog of protein complexes from the interaction screen data. An advantage of this model is that it places no prior constraints on the number of complexes and automatically infers the number of significant complexes from the data. Validation results using affinity purification/mass spectrometry experimental data from yeast RNA-processing complexes indicate that our method is capable of partitioning the data in a biologically meaningful way. A supplementary web site containing larger versions of the figures is available at http://public.kgi.edu/wild/PSBO6/index.html.

Algorithms↗

Performance of four ribosomal DNA regions to infer higher-level phylogenetic relationships of inoperculate euascomycetes (Leotiomyceta).

The inoperculate euascomycetes are filamentous fungi that form saprobic, parasitic, and symbiotic associations with a wide variety of animals, plants, cyanobacteria, and other fungi. The higher-level relationships of this economically important group have been unsettled for over 100 years. A data set of 55 species was assembled including sequence data from nuclear and mitochondrial small and large subunit rDNAs for each taxon; 83 new sequences were obtained for this study. Parsimony and Bayesian analyses were performed using the four-region data set and all 14 possible subpartitions of the data. The mitochondrial LSU rDNA was used for the first time in a higher-level phylogenetic study of ascomycetes and its use in concatenated analyses is supported. The classes that were recognized in Leotiomyceta (=inoperculate euascomycetes) in a classification by Eriksson and Winka [Myconet 1 (1997) 1] are strongly supported as monophyletic. The following classes formed strongly supported sister-groups: Arthoniomycetes and Dothideomycetes, Chaetothyriomycetes and Eurotiomycetes, and Leotiomycetes and Sordariomycetes. Nevertheless, the backbone of the euascomycete phylogeny remains poorly resolved. Bayesian posterior probabilities were always higher than maximum parsimony bootstrap values, but converged with an increase in gene partitions analyzed in concatenated analyses. Comparison of five recent higher-level phylogenetic studies in ascomycetes demonstrates a high degree of uncertainty in the relationships between classes.

Ascomycota↗

Inference and uncertainty in radiology.

This paper seeks to enhance understanding of the philosophical underpinnings of our discipline and the resulting practical implications. Radiology reports exist in order to convey new knowledge about a patient's condition based on empiric observations from anatomic or functional images of the body. The route to explanation and prediction from empiric evidence is mostly through inference based on inductive (and sometimes abductive) arguments. The conclusions of inductive arguments are, by definition, contingent and provisional. Therefore, it is necessary to deal in some way with the uncertainty of inferential conclusions (i.e. interpretations) made in radiology reports. Two paradigms for managing uncertainty in natural sciences exist in dialectic tension with each other. These are the frequentist and Bayesian theories of probability. Tension between them is mirrored during routine interactions among radiologists and clinicians. I will describe these core issues and argue that they are quite relevant to routine image interpretation and reporting.

Bayes Theorem↗

Effect of unsampled populations on the estimation of population sizes and migration rates between sampled populations.

Current estimators of gene flow come in two methods; those that estimate parameters assuming that the populations investigated are a small random sample of a large number of populations and those that assume that all populations were sampled. Maximum likelihood or Bayesian approaches that estimate the migration rates and population sizes directly using coalescent theory can easily accommodate datasets that contain a population that has no data, a so-called 'ghost' population. This manipulation allows us to explore the effects of missing populations on the estimation of population sizes and migration rates between two specific populations. The biases of the inferred population parameters depend on the magnitude of the migration rate from the unknown populations. The effects on the population sizes are larger than the effects on the migration rates. The more immigrants from the unknown populations that are arriving in the sample populations the larger the estimated population sizes. Taking into account a ghost population improves or at least does not harm the estimation of population sizes. Estimates of the scaled migration rate M (migration rate per generation divided by the mutation rate per generation) are fairly robust as long as migration rates from the unknown populations are not huge. The inclusion of a ghost population does not improve the estimation of the migration rate M; when the migration rates are estimated as the number of immigrants Nm then a ghost population improves the estimates because of its effect on population size estimation. It seems that for 'real world' analyses one should carefully choose which populations to sample, but there is no need to sample every population in the neighbourhood of a population of interest.

Bayes Theorem↗

A fully Bayesian model to cluster gene-expression profiles.

MOTIVATION: With cDNA or oligonucleotide chips, gene-expression levels of essentially all genes in a genome can be simultaneously monitored over a time-course or under different experimental conditions. After proper normalization of the data, genes are often classified into co-expressed classes (clusters) to identify subgroups of genes that share common regulatory elements, a common function or a common cellular origin. With most methods, e.g. k-means, the number of clusters needs to be specified in advance; results depend strongly on this choice. Even with likelihood-based methods, estimation of this number is difficult. Furthermore, missing values often cause problems and lead to the loss of data. RESULTS: We propose a fully probabilistic Bayesian model to cluster gene-expression profiles. The number of classes does not need to be specified in advance; instead it is adjusted dynamically using a Reversible Jump Markov Chain Monte Carlo sampler. Imputation of missing values is integrated into the model. With simulations, we determined the speed of convergence of the sampler as well as the accuracy of the inferred variables. Results were compared with the widely used k-means algorithm. With our method, biologically related co-expressed genes could be identified in a yeast transcriptome dataset, even when some values were missing. AVAILABILITY: The code is available at http://genome.tugraz.at/BayesianClustering/

Algorithms↗

On the change of support problem for spatio-temporal data.

In practice, spatial data are sometimes collected at points (i.e. point-referenced data) and at other times are associated with areal units (i.e. block data). The change of support problem is concerned with inference about the values of a variable at points or blocks different from those at which it has been observed. In the context of block data which can be sensibly viewed as averaging over point data, we propose a unifying approach for prediction from points to points, points to blocks, blocks to points, and blocks to blocks. The approach includes fully Bayesian kriging. We also extend our approach to the the case of spatio-temporal data, wherein a judicious specification of spatio-temporal association enables manageable computation. Exemplification of the static spatial case is provided using a dataset of point-level ozone measurements in the Atlanta, Georgia metropolitan area. The dynamic spatial case is illustrated using a temporally extended version of this dataset, enabling comparison at the common time point.*To whom correspondence should be addressed.

Journal Article↗

Mining Stored-Specimen Studies for Information about Cancer Natural History.

The advent of new multicancer early detection tests and publication of early diagnostic results have generated expectations of clinical benefit from multicancer screening. The clinical benefit of a cancer screening test depends critically on disease natural history, which is typically learned from prospective screening studies. Retrospective studies of stored blood specimens are important in learning about a test's preclinical diagnostic performance but have rarely been used to infer natural history. The extent to which these studies might be harnessed to also learn natural history is discussed in the context of an article in this issue that infers the combined natural history of a range of cancers targeted by a multicancer early detection test using a case-control subsample of specimens from a large cohort study. The critical question concerns the identifiability of key transition rates in multistate models of natural history alongside state-specific sensitivities. The article suggests that these parameters are estimable within a Bayesian framework that leverages prior information about test sensitivity from diagnostic studies. We offer a heuristic discussion of identifiability in this setting and encourage formal study to determine the extent to which models with varying degrees of complexity may be learned from stored-specimen studies. See related article by Dai et al., p. 1535.

Humans↗

Identification of a novel HIV-1 circulating recombinant form (CRF209_cpx) and its descendant unique recombinant form (URF) CRF209_cpx/B among MSM in Guangdong, southern China.

BACKGROUND: The epidemic of human immunodeficiency virus type 1 (HIV-1) continues to pose a significant global health challenge, with increasing genetic diversity. The co-circulation of multiple subtypes among the local population facilitates the emergence of unique or circulating recombinant forms (URFs or CRFs). In China, the predominant strains include CRF07_BC, CRF01_AE, CRF55_01B, and subtype B. This study characterizes a novel CRF209_cpx and its descendant recombinant CRF209_cpx/B among men who have sex with men (MSM) in Guangdong, southern China. METHODS: Individuals infected with URFs with similar genetic characteristics were recruited during routine surveillance of pretreatment drug resistance. Near full-length genomes (NFLGs) were amplified with two overlapping fragments using a serial dilution nested PCR approach after reverse transcription. We used SimPlot and IQ-TREE softwares to conduct recombination analyses and phylogenetic inferences. Time-scaled maximum clade credibility (MCC) phylogenetic trees were reconstructed using BEAST software to estimate evolutionary origins. Genotypic drug resistance mutations were interpreted via the Stanford HIV Database, and coreceptor usage was predicted using geno2pheno coreceptor 2.5 and the HIVcoPRED tool. RESULTS: Four NFLG sequences were obtained and identified as a novel CRF209_cpx, generated by recombination among CRF01_AE, CRF07_BC and subtype B. Phylogenetic analyses revealed that all the parental segments clustered with lineages prevalent among MSM in China. Bayesian evolutionary analysis estimated that the most recent common ancestor (tMRCA) of CRF209_cpx to have evolved between 2011 and 2013. The fifth strain was identified as a URF recombined from nascent CRF209_cpx and B. No transmitted drug resistance mutation was detected in these five sequences. The four CRF209_cpx sequences primarily utilized the CXCR4 coreceptor, while the URF exhibited R5/X4 dual tropism. CONCLUSIONS: The emergence of the complex CRF209_cpx and novel URF of CRF209_cpx/B highlights the active HIV-1 epidemic within the MSM population in Guangdong, underscoring the necessity for enhanced molecular surveillance and precise public health intervention in this key population.

HIV-1↗

Borrowing strength from external trials in a meta-analysis.

There exists a variety of situations in which a random effects meta-analysis might be undertaken using a small number of clinical trials. A problem associated with small meta-analyses is estimating the heterogeneity between trials. To overcome this problem, information from other related studies may be incorporated into the meta-analysis. A Bayesian approach to this problem is presented using data from previous meta-analyses in the same therapeutic area to formulate a prior distribution for the heterogeneity. The treatment difference parameters are given non-informative priors. Further, related trials which compare one or other of the treatments of interest with a common third treatment are included in the model to improve inference on both the heterogeneity and the treatment difference. Two approaches to estimating relative efficacy are considered, namely a general parametric approach and a method explicit to binary data. The methodology is illustrated using data from 26 clinical trials which investigate the prevention of cirrhosis using beta-blockers and sclerotherapy. Both sources of external information lead to more precise posterior distributions for all parameters, in particular that representing heterogeneity.

Adrenergic beta-Antagonists↗