Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Reasoning requirements for diagnosis of heart disease.

Over the past dozen years, the Heart Disease Program (HDP) has been developed to assist physicians in reasoning about cardiovascular disorders. Driven by several evaluations, the inference mechanism has progressed from a logic based model, to a Bayesian Probability Network (BPN) and finally a pseudo-Bayesian network with temporal and severity reasoning. Though aspects of cardiovascular reasoning are handled well by BPNs, temporal reasoning, homeostatic feedback mechanisms and effects of disease severities require additional inference strategies. This article discusses how these reasoning problems are handled, and deals with closely linked issues in building the user interface to collect detailed cardiovascular data and provide clear explanations of diagnoses.

Artificial Intelligence↗

Meta-analysis to determine the incidence of obstetric anal sphincter damage.

BACKGROUND: The reported incidence of anal sphincter injury after first (11.5-35.0 per cent) and subsequent (3.4-12.1 per cent) vaginal deliveries varies widely. In addition, the reported incidence of associated faecal incontinence ranges from zero to 68.2 per cent. The aim of this study was to perform a meta-analysis of reported incidences of postpartum anal sphincter defect diagnosed by endoanal ultrasonography (EAUS) and associated incidences of faecal incontinence. METHODS: A Medline search yielded five studies with more than 100 subjects who underwent EAUS after childbirth for evaluation of anal sphincter disruption and who were questioned about symptoms of faecal incontinence, defined as any impairment in flatus and stool control but not including urgency of defaecation. A Bayesian meta-analysis was performed to produce one inference while accounting for potential heterogeneity among the five study populations. RESULTS: Meta-analysis of 717 vaginal deliveries revealed a 26.9 per cent incidence of anal sphincter defect in primiparous women and an 8.5 per cent incidence of new sphincter defects in multiparous women. Overall, 29.7 per cent of anal sphincter defects were symptomatic. Some 3.4 per cent of women experienced postpartum faecal incontinence without an anal sphincter defect. In a Bayesian calculation, the probability of postpartum faecal incontinence due to a sphincter defect was 76.8-82.8 per cent. CONCLUSION: : The incidence of occult anal sphincter disruption following vaginal delivery is much higher than commonly estimated. However, at least two-thirds of occult defects are asymptomatic postpartum. The probability of faecal incontinence associated with an anal sphincter defect was 76.8-82.8 per cent.

Anal Canal↗

Bayesian analysis of geographical variation in the incidence of Type I diabetes in Finland.

AIMS/HYPOTHESIS: In Finland, the incidence of Type I (insulin-dependent) diabetes mellitus among children aged 14 years or under is the highest in the world. The increase in incidence is approximately 3% per year. A marked geographical variation in incidence was reported in Finland during the late 1980s. Our aim was to explore the most recent regional pattern in incidence of Type I diabetes in Finland. METHODS: Data on the nationwide incidence of childhood diabetes in Finland was obtained from the Prospective Childhood Diabetes Registry for the periods 1987-1991 and 1992-1996. Population data was obtained from the National Population Registry. The geographical pattern of incidence was studied applying a Bayesian hierarchical approach and Geographical Information Systems. The inferences from the data was based on the estimated geographical intensity of diabetes. RESULTS: There was a clear evidence of geographic variation for the risk of childhood diabetes during the entire 10-year period. The high-risk areas were found in the wide belt crossing the central part of Finland. Comparison of the estimated intensity of diabetes between the two 5-year periods showed that the geographical pattern of diabetes risk has changed over time. Our analyses also confirmed the existence of a few persistent high-risk and low-risk areas in Finland. CONCLUSION/INTERPRETATION: The finding of high-risk areas of childhood Type I diabetes suggests that specific genetic or environmental risk factors have become greater in certain geographic locations in Finland.

Adolescent↗

Major gene detection for fusiform rust resistance using Bayesian complex segregation analysis in loblolly pine.

The presence of major genes affecting rust resistance of loblolly pine was investigated in a progeny population that was generated with a half-diallel mating of six parents. A Bayesian complex segregation analysis was used to make inference about a mixed inheritance model (MIM) that included polygenic effects and a single major gene effect. Marginalizations were achieved by using Gibbs sampler. A parent block sampling by which genotypes of a parent and its offspring were sampled jointly was implemented to improve mixing. The MIM was compared with a pure polygenic model (PM) using Bayes factor. Results showed that the MIM was a better model to explain the inheritance of rust resistance than the pure PM in the diallel population. A large major gene variance component estimate (> 50% of total variance), indicated the existence of major genes for rust resistance in the studied loblolly pine population. Based on estimations of parental genotypes, it appears that there may be two or more major genes affecting disease phenotypes in this diallel population.

Basidiomycota↗

The evaluation of evidence in the forensic investigation of fire incidents (Part I): an approach using Bayesian networks.

The forensic investigation of the origin and cause of a fire incident is a particularly demanding area of expertise. As the available evidence is often incomplete or vague, uncertainty is a key element. The present study is an attempt to approach this through the use of Bayesian networks, which have been found useful in assisting human reasoning in a variety of disciplines in which uncertainty plays a central role. The present paper describes the construction of a Bayesian network (BN) and its use for drawing inferences about propositions of interest, based upon a single, possibly non replicable item of evidence: detected residual quantities of a flammable liquid in fire debris.

Journal Article↗

Phylogenetic relationships between members of the crucifer pathogenic Leptosphaeria maculans species complex as shown by mating type (MAT1-2), actin, and beta-tubulin sequences.

The dothideomycetous fungus Leptosphaeria maculans comprises a complex of species differing in specificity and pathogenicity on Brassica napus. Twenty-eight isolates were investigated and compared to 20 other species of the Pleosporales order. Sequences of the mating type MAT1-2 (23), fragments of actin (48) and beta-tubulin (45) genes were determined and used for phylogenetic analyses inferred by maximum parsimony, distance, maximum likelihood, and Bayesian approaches. These different approaches using single genes essentially confirmed findings using concatanated sequences. L. maculans formed a monophyletic group separate from Leptosphaeria biglobosa. The L. biglobosa clade encompasses five sub-clades; this finding is consistent with classification made previously on the basis of internal-transcribed sequences of the ribosomal DNA repeat. The propensity for purifying and neutral evolution of the three genes was determined using sliding window analysis, a technique not previously applied to genes of filamentous fungi. For members of the L. maculans species complex, this approach showed that in comparison to actin and beta-tubulin, exonic sequences of MAT1-2 were more diverse and appeared to evolve at a faster rate. However, different regions of MAT1-2 displayed different degrees of sequence conservation. The more conserved upstream region (including the High Mobility Group domain) may be better suited for interspecies differentiation, while the more diverse downstream region is more appropriate for intraspecies comparisons.

Actins↗

Evaluation of scientific evidence using Bayesian networks.

Bayesian networks provide a valuable aid for representing epistemic relationships in a body of uncertain evidence. The paper proposes some simple Bayesian networks for standard analysis of patterns of inference concerning scientific evidence, with a discussion of the rationale behind the nets, the corresponding probabilistic formulas, and the required probability assessments.

Adult↗

Estimating the distribution of worm burden and egg excretion of Schistosoma japonicum by risk group in Sichuan Province, China.

During autumn 2000 an extensive cross-sectional survey of the prevalence of Schistosomiasis japonicum was conducted among about 4000 villagers within 20 villages in the Anning River Valley located in the southwestern Sichuan Province. Two procedures were used to assess infection status, the Kato-Katz thick smear procedure and a miracidia hatch test. Whereas the Kato-Katz procedure provides information on both prevalence and intensity, the hatch test provides only prevalence data, albeit on a much larger volume of stool. In addition, we performed Kato-Katz smears for 15 consecutive samples on a subset of 15 individuals. The proportion of both hatch-test and Kato-Katz positive individuals in the larger cross-sectional survey was 25%. The goal of the study was to estimate both the egg and worm distributions among risk groups using both the hatch and Kato-Katz tests from the cross-sectional data and the repeated Kato-Katz smears from the longitudinal data sets. As a prelude to parameter estimation, individuals were classified into risk groups by natural village and occupation; the proportion of Kato-Katz positive subjects among the risk groups varied from 10% to 60%. We used the statistical model of de Vlas et al. (1992) and Bayesian techniques to derive both estimates of and inference about the worm and egg distribution parameters. The parameter estimates imply (1) similar eggs per gram stool (e.p.g.) per worm pair compared with earlier estimates, (2) a range of worm burdens among the risk groups and (3) estimates of risk heterogeneity within groups is sensitive to prior information on the within-person variability in egg excretion.

Age Factors↗

Detecting recombination with MCMC.

MOTIVATION: We present a statistical method for detecting recombination, whose objective is to accurately locate the recombinant breakpoints in DNA sequence alignments of small numbers of taxa (4 or 5). Our approach explicitly models the sequence of phylogenetic tree topologies along a multiple sequence alignment. Inference under this model is done in a Bayesian way, using Markov chain Monte Carlo (MCMC). The algorithm returns the site-dependent posterior probability of each tree topology, which is used for detecting recombinant regions and locating their breakpoints. RESULTS: The method was tested on a synthetic and three real DNA sequence alignments, where it was found to outperform the established detection methods PLATO, RECPARS, and TOPAL.

Algorithms↗

Discriminating between rate heterogeneity and interspecific recombination in DNA sequence alignments with phylogenetic factorial hidden Markov models.

MOTIVATION: A recently proposed method for detecting recombination in DNA sequence alignments is based on the combination of hidden Markov models (HMMs) with phylogenetic trees. Although this method was found to detect breakpoints of recombinant regions more accurately than most existing techniques, it inherently fails to distinguish between recombination and rate variation. In the present paper, we propose to marry the phylogenetic tree to a factorial HMM (FHMM). The states of the first hidden chain represent tree topologies, whereas the states of the second independent hidden chain represent different global scaling factors of the branch lengths. Inference is done in terms of a hierarchical Bayesian model, where parameters and hidden states are sampled from the posterior distribution with Gibbs sampling. RESULTS: We have tested the proposed model on various synthetic and real-world DNA sequence alignments. The simulation results suggest that as opposed to the standard phylogenetic HMM, the phylogenetic FHMM clearly distinguishes between recombination and rate heterogeneity and thereby avoids the prediction of spurious recombinant regions. AVAILABILITY: The proposed method has been implemented in a MATLAB package that extends Kevin Murphy's HMM toolbox. Software and data used in our study are available from http://www.bioss.sari.ac.uk/~dirk/Supplements

Algorithms↗

Time dependency of molecular rate estimates and systematic overestimation of recent divergence times.

Studies of molecular evolutionary rates have yielded a wide range of rate estimates for various genes and taxa. Recent studies based on population-level and pedigree data have produced remarkably high estimates of mutation rate, which strongly contrast with substitution rates inferred in phylogenetic (species-level) studies. Using Bayesian analysis with a relaxed-clock model, we estimated rates for three groups of mitochondrial data: avian protein-coding genes, primate protein-coding genes, and primate d-loop sequences. In all three cases, we found a measurable transition between the high, short-term (< 1-2 Myr) mutation rate and the low, long-term substitution rate. The relationship between the age of the calibration and the rate of change can be described by a vertically translated exponential decay curve, which may be used for correcting molecular date estimates. The phylogenetic substitution rates in mitochondria are approximately 0.5% per million years for avian protein-coding sequences and 1.5% per million years for primate protein-coding and d-loop sequences. Further analyses showed that purifying selection offers the most convincing explanation for the observed relationship between the estimated rate and the depth of the calibration. We rule out the possibility that it is a spurious result arising from sequence errors, and find it unlikely that the apparent decline in rates over time is caused by mutational saturation. Using a rate curve estimated from the d-loop data, several dates for last common ancestors were calculated: modern humans and Neandertals (354 ka; 222-705 ka), Neandertals (108 ka; 70-156 ka), and modern humans (76 ka; 47-110 ka). If the rate curve for a particular taxonomic group can be accurately estimated, it can be a useful tool for correcting divergence date estimates by taking the rate decay into account. Our results show that it is invalid to extrapolate molecular rates of change across different evolutionary timescales, which has important consequences for studies of populations, domestication, conservation genetics, and human evolution.

Animals↗

Risk of small-for-gestational age is associated with common anti-inflammatory cytokine polymorphisms.

BACKGROUND: Anti-inflammatory cytokines play a key role in pregnancy maintenance. Genetic variation in anti-inflammatory cytokines could influence a woman's risk of adverse reproductive outcomes. METHODS: We investigated the relationship of polymorphisms in interleukin 4 (IL4), IL5, IL10, IL13, and transforming growth factor (TGFbeta1) with spontaneous preterm birth and small-for-gestational age (SGA) in a nested case-control study of a prospective pregnancy cohort. Women were recruited between 24 and 29 weeks' gestation at the Wake County and University of North Carolina, Chapel Hill obstetric clinics between February 1996 and June 2000. We inferred haplotypes using the EM algorithm and the Bayesian method, PHASE. Semi-Bayesian hierarchical logistic regression was used to obtain odds ratio (OR) estimates and 95% confidence intervals (CIs) for each polymorphism. RESULTS: African-American mothers who carried the IL4 GCC haplotype had greater risk of spontaneous preterm birth (OR = 2.9; 95% CI = 1.2-7.4). In white mothers, carriers of the "low-producing" IL4 CC and IL10 ATA haplotypes had markedly reduced risk of SGA (for the CC haplotype, 0.2 [0.0-1.2]; for the ATA haplotype, 0.5 [0.3-0.8]), whereas carriers of the "high-producing" IL4(-589)T variant had increased risk of SGA in both African-American and white mothers. CONCLUSIONS: Variants related to decreased anti-inflammatory cytokine production may lower risk of SGA. Furthermore, the same mechanism that protects against SGA might increase risk of spontaneous preterm birth.

Black or African American↗

Risk of spontaneous preterm birth is associated with common proinflammatory cytokine polymorphisms.

BACKGROUND: Preliminary data suggest that common genetic variation in immune response genes can contribute to the risk for spontaneous preterm birth and possibly small-for-gestational age (SGA). METHODS: We investigated the relationship of polymorphisms in 6 cytokine genes associated with inflammation-interleukin (IL)1alpha, IL1beta, IL2, IL6, tumor necrosis factor (TNF), and lymphotoxin alpha (LTA)-with spontaneous preterm and SGA birth in a nested case-control study drawn from a prospective pregnancy cohort. Women were recruited between 24 and 29 weeks' gestation at the Wake County and University of North Carolina, Chapel Hill obstetric clinics between February 1996 and June 2000. We inferred haplotypes using the EM algorithm and the Bayesian method, PHASE. We then compared haplotype frequency distributions and implemented semi-Bayesian hierarchical logistic regression analyses to obtain odds ratio (OR) estimates and 95% confidence intervals (CIs) for each polymorphism. RESULTS: Two haplotypes spanning the TNF/LTA genes were associated with increased risk for spontaneous preterm birth in white subjects (for the AGG haplotype, OR = 1.5 [95% CI=0.8-2.6]; for the GAC haplotype, 1.6 [0.9-2.9]). Additionally, carriers of the GAG haplotype were found to have decreased risk of spontaneous preterm birth (0.6; 0.3-1.0). The TNF(-488)A and LTA(IVS1-82)C variants, constituents of the AGG and GAC haplotypes respectively, were also strongly associated with increased risk of spontaneous preterm birth. CONCLUSIONS: Our results suggest that common genetic variants in proinflammatory cytokine genes could influence the risk for spontaneous preterm birth. Selected TNF/LTA haplotypes were associated with spontaneous preterm birth in both African-American and white subjects. Our data do not support an inflammatory etiology for SGA.

Black or African American↗

Dispersal, vicariance, and timing of diversification in Nothonotus darters.

The species diversity of North American freshwater fishes is unparalleled among temperate regions of the planet. This diversity is concentrated in the Central Highlands of eastern North America and this distribution pattern has inspired different models involving either dispersal or vicariance to explain the high species diversity of North American fishes. The most popular of these models is the Central Highlands vicariance hypothesis (CHVH), which proposes an ancient and diverse widespread fauna that existed across a previously continuous highland landscape that is much different from today. The mechanisms of isolation in the CHVH involve specific instances of vicariance that affected several diverse lineages of Central Highlands fishes. We tested predictions of the CHVH and alternative models using a cytochrome b-inferred phylogeny of the darter clade Nothonotus. A Bayesian mixed-model method was used for phylogenetic analysis. The phylogenetic data set included all 20 recognized Nothonotus species, and most species were represented with multiple sequences. We were able to convert genetic branch lengths to absolute age using external fossil calibrations in the freshwater perciform fish clade Centrarchidae. Using a well-resolved Nothonotus phylogeny and divergence time estimates, we identify equal numbers of instances of both vicariance and dispersal among disjunct regions of the Central Highlands, biogeographic pseudocongruence, rather recent speciation in Nothonotus, and a surprisingly large amount of speciation within highland areas. With regard to Nothonotus, previous Central Highlands biogeographic models offer little in the way of providing possible mechanisms responsible for diversification in the clade. Patterns of speciation in Nothonotus are similar to those discovered in recent efforts that have included speciation as a parameter into classic models of island biogeography.

Animals↗

Protein identification by tandem mass spectrometry and sequence database searching.

The shotgun proteomics strategy, based on digesting proteins into peptides and sequencing them using tandem mass spectrometry (MS/MS), has become widely adopted. The identification of peptides from acquired MS/MS spectra is most often performed using the database search approach. We provide a detailed description of the peptide identification process and review the most commonly used database search programs. The appropriate choice of the search parameters and the sequence database are important for successful application of this method, and we provide general guidelines for carrying out efficient analysis of MS/MS data. We also discuss various reasons why database search tools fail to assign the correct sequence to many MS/MS spectra, and draw attention to the problem of false-positive identifications that can significantly diminish the value of published data. To assist in the evaluation of peptide assignments to MS/MS spectra, we review the scoring schemes implemented in most frequently used database search tools. We also describe statistical approaches and computational tools for validating peptide assignments to MS/MS spectra, including the concept of expectation values, reversed database searching, and the empirical Bayesian analysis of PeptideProphet. Finally, the process of inferring the identities of the sample proteins given the list of peptide identifications is outlined, and the limitations of shotgun proteomics with regard to discrimination between protein isoforms are discussed.

Amino Acid Sequence↗

Phylogeographic history and gene flow among giant Galápagos tortoises on southern Isabela Island.

Volcanic islands represent excellent models with which to study the effect of vicariance on colonization and dispersal, particularly when the evolution of genetic diversity mirrors the sequence of geological events that led to island formation. Phylogeographic inference, however, can be particularly challenging for recent dispersal events within islands, where the antagonistic effects of land bridge formation and vicariance can affect movements of organisms with limited dispersal ability. We investigated levels of genetic divergence and recovered signatures of dispersal events for 631 Galápagos giant tortoises across the volcanoes of Sierra Negra and Cerro Azul on the island of Isabela. These volcanoes are among the most recent formations in the Galápagos (<0.7 million years), and previous studies based on genetic and morphological data could not recover a consistent pattern of lineage sorting. We integrated nested clade analysis of mitochondrial DNA control region sequences, to infer historical patterns of colonization, and a novel Bayesian multilocus genotyping method for recovering evidence of recent migration across volcanoes using eleven microsatellite loci. These genetic studies illuminate taxonomic distinctions as well as provide guidance to possible repatriation programs aimed at countering the rapid population declines of these spectacular animals.

Animals↗

[Statistical models for spatial analysis in parasitology].

The simplest way to study the spatial pattern of a disease is the geographical representation of its cases (or some indicators of them) over a map. Maps based on raw data are generally "wrong" since they do not take into consideration for sampling errors. Indeed, the observed differences between areas (or points in the map) are not directly interpretable, as they derive from the composition of true, structural differences and of the noise deriving from the sampling process. This problem is well known in human epidemiology, and several solutions have been proposed to filter the signal from the noise. These statistical methods are usually referred to as Disease Mapping. In geographical analysis a first goal is to evaluate the statistical significance of the heterogeneity between areas (or points). If the test indicates rejection of the hypothesis of homogeneity the following task is to study the spatial pattern of the disease. The spatial variability of risk is usually decomposed into two terms: a spatially structured (clustering) and a non spatially structured (heterogeneity) one. The heterogeneity term reflects spatial variability due to intrinsic characteristics of the sampling units (e.g. igienic conditions of farms), while the clustering term models the association due to proximity between sampling units, that usually depends on ecological conditions that vary over the study area and that affect in similar way breedings that are close to each other. Hierarchical bayesian models are the main tool to make inference over the clustering and heterogeneity components. The results are based on the marginal posterior distributions of the parameters of the model, that are approximated by Monte Carlo Markov Chain methods. Different models can be defined depending on the terms that are considered, namely a model with only the clustering term, a model with only the heterogeneity term and a model where both are included. Model selection criteria based on a compromise between degree of complexity and goodness of fit are then needed to discriminate among them, because each specification has a different biological meaning. Our aim is to demonstrate that these techniques can be used to study the geographical distribution of a parasite infection. Our analyses are based on data collected in 142 farms of the province of Latina. In each breeding a fixed number of sheeps has been sampled (20) and checked for the presence of C. daubneyi. We have specified a Binomial model for the proportion of infected animals in each breeding. The heterogeneity component is modelled in a standard way, while we have used different prior specifications for the clustering term to show how they affect the results. When we use the usual specification also for clustering, the two models show a completely different spatial pattern of infection, probably because the intrinsic spatial structure of the clustering term tend to bias our inferences. The selection criterion indicates in this case the heterogeneity model as the "best" one. However, if we modify the prior so that a lower degree of spatial interaction is assumed, the clustering model is less complex and its goodness of fit better and it should be preferred.

Animal Husbandry↗

Statistical significance--a misconstrued notion in medical research.

The P-value is the significance probability of obtaining a value of the test statistic that is as extreme, in relation to the null hypothesis, as that observed. Medical researchers may, in some situations, disagree on its appropriate use or on its interpretation as a summary measure of consistency with the null hypothesis in a particular data set. More informative statistical measures such as the likelihood ratio and the Bayesian posterior probability have been suggested for drawing inferences from clinical trials and epidemiologic studies. Causal inference is not statistical in nature; rather it strives to provide scientific explanations or criticisms of proposed explanations that would describe the observed data pattern. In this context, it is important to remember that a finding may not be medically important, or a causal hypothesis may even not be true even if a study shows a significant P-value.

Bayes Theorem↗