Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Estimating biomarker-based HIV incidence using prevalence data in high risk groups with missing outcomes.

The novel two-step serologic sensitive/less sensitive testing algorithm for detecting recent HIV seroconversion (STARHS) provides a simple and practical method to estimate HIV-1 incidence using cross-sectional HIV seroprevalence data. STARHS has been used increasingly in epidemiologic studies. However, the uncertainty of incidence estimates using this algorithm has not been well described, especially for high risk groups or when missing data is present because a fraction of sensitive enzyme immunoassay (EIA) positive specimens are not tested by the less sensitive EIA. Ad hoc methods used in practice provide incorrect confidence limits and thus may jeopardize statistical inference. In this report, we propose maximum likelihood and Bayesian methods for correctly estimating the uncertainty in incidence estimates obtained using prevalence data with a fraction missing, and extend the methods to regression settings. Using a study of injection drug users participating in a drug detoxification program in New York city as an example, we demonstrated the impact of underestimating the uncertainty in incidence estimates using ad hoc methods. Our methods can be applied to estimate the incidence of other diseases from prevalence data using similar testing algorithms when missing data is present.

Bayes Theorem↗

Linkage disequilibrium assessment via log-linear modeling of SNP haplotype frequencies.

Analyses of high-density single-nucleotide polymorphism (SNP) data, such as genetic mapping and linkage disequilibrium (LD) studies, require phase-known haplotypes to allow for the correlation between tightly linked loci. However, current SNP genotyping technology cannot determine phase, which must be inferred statistically. In this paper, we present a new Bayesian Markov chain Monte Carlo (MCMC) algorithm for population haplotype frequency estimation, particularly in the context of LD assessment. The novel feature of the method is the incorporation of a log-linear prior model for population haplotype frequencies. We present simulations to suggest that 1) the log-linear prior model is more appropriate than the standard coalescent process in the presence of recombination (>0.02 cM between adjacent loci), and 2) there is substantial inflation in measures of LD obtained by a "two-stage" approach to the analysis by treating the "best" haplotype configuration as correct, without regard to uncertainty in the recombination process.

Algorithms↗

Molecular phylogeny and biogeography of an ancient Holarctic lineage of mygalomorph spiders (Araneae: Antrodiaetidae: Antrodiaetus).

The mygalomorph spider genera Antrodiaetus and Atypoides (Antrodiaetidae) belong to an ancient lineage that has persisted since at least the Cretaceous. These spiders display a classic disjunct Holarctic distribution with species in the eastern Palaearctic plus the western and eastern Nearctic. Prior phylogenetic analyses of this group have been proposed on the basis of morphology, but lack strong support and independent corroboration. Here we present the first phylogenetic analysis of species-level relationships based on molecular data obtained from the mitochondrial (cytochrome c oxidase subunit I) and nuclear (18S and 28S rRNA) genomes. Analyses corroborate earlier findings that Atypoides forms a paraphyletic grade with respect to Antrodiaetus, and consequently, that genus is formally synonymized under Antrodiaetus. In addition, our results support the relatively early divergence of Antrodiaetus roretzi. Antrodiaetus pacificus is "paraphyletic" with respect to the A. lincolnianus group and is likely an assemblage of numerous species. The final topology based on a combined molecular dataset, in conjunction with two different molecular dating techniques (penalized likelihood plus a Bayesian approach) and ancestral distribution reconstructions, was used to infer the historical biogeography of these spiders. Trans-Beringian and trans-Atlantic routes appear to account for the present-day distribution of Antrodiaetus in Japan and North America. Future studies on Antrodiaetus phylogeny will be used to address questions regarding morphological stasis and the evolution of quantitative morphological characters.

Animals↗

Evolutionary changes of metabolic networks and their biosynthetic capacities.

The metabolic networks of different species show a large variety in their structural design. In this work, the evolution of functional properties of metabolism in relation with metabolic network structure is investigated. The metabolism of ancestral species is inferred from the metabolism of contemporary species using a Bayesian network model for metabolism evolution. Subsequently, these networks are analysed with the recently developed method of network expansion. This method allows for a structural analysis of metabolic networks as well as a quantification of network functions in terms of their synthesising capacities when they are provided with certain external resources. The evolutionary dynamics of one particular network function: the metabolic expansion of glucose is investigated.

Adaptation, Physiological↗

Recombination hotspots as a point process.

The variation of the recombination rate along chromosomal DNA is one of the important determinants of the patterns of linkage disequilibrium. A number of inferential methods have been developed which estimate the recombination rate and its variation from population genetic data. The majority of these methods are based on modelling the genealogical process underlying a sample of DNA sequences and thus explicitly include a model of the demographic process. Here we propose a different inferential procedure based on a previously introduced framework where recombination is modelled as a point process along a DNA sequence. The approach infers regions containing putative hotspots based on the inferred minimum number of recombination events; it thus depends only indirectly on the underlying population demography. A Poisson point process model with local rates is then used to infer patterns of recombination rate estimation in a fully Bayesian framework. We illustrate this new approach by applying it to several population genetic datasets, including a region with an experimentally confirmed recombination hotspot.

Bayes Theorem↗

Bayesian modeling of incidence and progression of disease from cross-sectional data.

In the absence of longitudinal data, the current presence and severity of disease can be measured for a sample of individuals to investigate factors related to disease incidence and progression. In this article, Bayesian discrete-time stochastic models are developed for inference from cross-sectional data consisting of the age at first diagnosis, the current presence of disease, and one or more surrogates of disease severity. Semiparametric models are used for the age-specific hazards of onset and diagnosis, and a normal underlying variable approach is proposed for modeling of changes with latency time in disease severity. The model accommodates multiple surrogates of disease severity having different measurement scales and heterogeneity among individuals in disease progression. A Markov chain Monte Carlo algorithm is described for posterior computation, and the methods are applied to data from a study of uterine leiomyoma.

Age of Onset↗

Evaluation of a new antibody-based enzyme-linked immunosorbent assay for the detection of bovine leukemia virus infection in dairy cattle.

The objective of this study was to validate a new blocking enzyme-linked immunosorbent assay (ELISA) (designated M108 for milk and S108 for serum samples) for detecting bovine leukemia virus (BLV) infection in dairy cattle. Milk, serum, and ethylenediaminetetraacetic acid-blood samples were collected from 524 adult Holstein cows originating from 6 dairy herds in Central Argentina. The M108 and S108 were compared with agar gel immunodiffusion (AGID), polymerase chain reaction and a commercial ELISA. Because there is currently no reference test capable of serving as a gold standard, the test sensitivity (SE) and specificity (SP) were evaluated by the use of a latent class model. Statistical inference was performed by classical maximum likelihood and by Bayesian techniques. The maximum-likelihood analysis was performed assuming conditional independence of tests, whereas the Bayesian approach allowed for conditional dependence. No clear conclusion could be drawn about conditional dependence of tests. Results with maximum likelihood (under conditional independence) and posterior Bayes (under conditional dependence) were practically the same. Conservative estimates of SE and SP (with 95% confidence intervals) for M108 were 98.6 (96.7; 99.6) and 96.7 (92.9; 98.8) and for S108 99.5 (98.2; 99.9) and 95.4 (90.9; 98.1), respectively. The ELISA 108 using either milk or serum to detect BLV-infected animals had comparable SE and SP with the official AGID and a commercial ELISA test, which are currently the most widely accepted tests for the serological diagnosis of BLV infection. Therefore, ELISA 108 can be used as an alternative test in monitoring and control programs.

Animals↗

Random regression test-day models with residuals following a Student's-t distribution.

First-lactation milk yield test-day records of Canadian Holsteins were analyzed by single-trait random regression test-day models that assumed normal or Student's-t distribution for residuals. Objectives were to test the performance of the robust statistical models that use heavy-tailed distributions for the residual effect. Models fitted were: Gaussian, Student's-t, and Student's-t with fixed number of degrees of freedom (equal to 5, 15, 30, 100 or 1000) for the t distribution. Bayesian methods with Gibbs sampling were used to make inferences about overall model plausibility through Bayes factors, posterior means for covariance components, estimated breeding values for regression coefficients, solutions for permanent environmental regressions, and residuals of the models. Bayes factors favored Student's-t model with the posterior mean of degrees of freedom equal to 2.4 over all other models, indicating very strong departure from normality. Number of outliers in Student's-t model was reduced by 35% in comparison with the Gaussian model. Differences in covariance components for regression coefficients between models were small, and rankings of animals based on additive genetic merit for the first two regression coefficients (total yield and persistency) were similar. Results from the Gaussian and Student's-t models with fixed degrees of freedom become more alike (smaller departures from normality for Student's-t models) with increasing number of degrees of freedom for the t-distributions. For any pair of Student's-t models, the one with the smaller number of degrees of freedom for the t-distribution was shown to be superior. Similarly, number of outliers increased with increasing degrees of freedom for the t distribution.

Animals↗

Prior (co)variances can improve multiple-trait across-country evaluations of weakly linked bull populations.

National genetic evaluation results for fore udder attachment from 9 Ayrshire populations were used to assess the impact of different uses of prior genetic correlations in multiple-trait across-country evaluations (MACE) on predicted international genetic merit. These Ayrshire populations were poorly connected; that is, 2% of the bulls had evaluations in 2 or more countries. Genetic correlations from the Holstein populations in the same countries were used as prior information to improve inferences of location parameters and international genetic merits. Fully Bayesian analyses using Gibbs sampling and computationally less demanding traditional MACE assuming a weighted average of prior and estimated Ayrshire genetic correlations were compared for 3 different prior degrees of belief and for different groups of bulls. Posterior means of genetic correlations estimated by Gibbs sampling were on average higher (+0.2) than those estimated by REML. Posterior heritabilities differed up to 0.2 units from those assumed in national genetic evaluations. Predicted genetic merit and international sire rankings of bulls with daughter information in the country of interest were not affected substantially by method of analysis and even less by varying prior degree of belief. Method of analysis had a larger impact on predicted genetic merit for bulls without daughter information in the country of interest. Here the average correlation between predicted genetic merit in different analyses ranged from 0.62 to 0.99. The predictive ability for young and randomly chosen bulls favored Bayesian MACE. The prior degree of belief did not have much impact on sire rankings and predictive ability, but intermediate prior degree of belief tended to perform best. All MACE analyses yielded nearly unbiased predictions. Traditional MACE assuming a simple weighted average of prior and estimated Ayrshire genetic correlations has been implemented by Interbull for routine international genetic evaluations.

Analysis of Variance↗

Bayesian approaches to modeling the conditional dependence between multiple diagnostic tests.

Many analyses of results from multiple diagnostic tests assume the tests are statistically independent conditional on the true disease status of the subject. This assumption may be violated in practice, especially in situations where none of the tests is a perfectly accurate gold standard. Classical inference for models accounting for the conditional dependence between tests requires that results from at least four different tests be used in order to obtain an identifiable solution, but it is not always feasible to have results from this many tests. We use a Bayesian approach to draw inferences about the disease prevalence and test properties while adjusting for the possibility of conditional dependence between tests, particularly when we have only two tests. We propose both fixed and random effects models. Since with fewer than four tests the problem is nonidentifiable, the posterior distributions are strongly dependent on the prior information about the test properties and the disease prevalence, even with large sample sizes. If the degree of correlation between the tests is known a priori with high precision, then our methods adjust for the dependence between the tests. Otherwise, our methods provide adjusted inferences that incorporate all of the uncertainty inherent in the problem, typically resulting in wider interval estimates. We illustrate our methods using data from a study on the prevalence of Strongyloides infection among Cambodian refugees to Canada.

Bayes Theorem↗

Can cognitive processes be inferred from neuroimaging data?

There is much interest currently in using functional neuroimaging techniques to understand better the nature of cognition. One particular practice that has become common is 'reverse inference', by which the engagement of a particular cognitive process is inferred from the activation of a particular brain region. Such inferences are not deductively valid, but can still provide some information. Using a Bayesian analysis of the BrainMap neuroimaging database, I characterize the amount of additional evidence in favor of the engagement of a cognitive process that can be offered by a reverse inference. Its usefulness is particularly limited by the selectivity of activation in the region of interest. I argue that cognitive neuroscientists should be circumspect in the use of reverse inference, particularly when selectivity of the region in question cannot be established or is known to be weak.

Brain↗

Tumour matrilysin expression predicts metastatic potential of stage I (pT1) colon and rectal cancers.

BACKGROUND AND AIMS: Nodal metastases are indisputable determinants of prognosis for colon and rectal cancer. Using classical histological criteria, many attempts to predict nodal metastasis have failed, preventing the adequate management of stage I (pT1) cancer. We investigated the role of tumour matrilysin in predicting metastatic potential, and discuss its potential use in individualising treatment of pT1 colon and rectal cancer. METHODS: The gene signature associated with nodal metastasis was investigated by cDNA array in 24 colon and rectal cancers. We studied 494 colon and rectal cancer patients to identify risk factors for nodal metastasis and evaluated the potential to predict nodal metastasis by either the logistic regression model or the Bayesian neural network model with built-in matrilysin. We then inferred possible causality of nodal metastasis from structural equation modelling. RESULTS: cDNA array revealed that matrilysin was maximally upregulated in the metastasis signature identified. Tumour matrilysin expression emerged as a stage independent risk factor for nodal metastasis, resulting in a similar predictive performance in receiver operating characteristic curve analysis in the two models. A Bayesian approach called automatic relevance determination identified matrilysin as one of the most relevant predictors examined. Structural equation modelling suggested possible direct causality between matrilysin and nodal metastasis. CONCLUSIONS: We have provided evidence that tumour matrilysin expression is a promising biomarker predicting nodal metastasis of colon and rectal cancer. Analysis of tumour matrilysin expression would help clinicians achieve the goal of individualised cancer treatment based on the metastatic potential of pT1 colon and rectal cancer.

Adenocarcinoma↗

Bayesian methods for a growth-curve degradation model with repeated measures.

The increasing reliability of some manufactured products has led to fewer observed failures in reliability testing. Thus, useful inference on the distribution of failure times is often not possible using traditional survival analysis methods. Partly as a result of this difficulty, there has been increasing interest in inference from degradation measurements made on products prior to failure. In the degradation literature inference is commonly based on large-sample theory and, if the degradation path model is nonlinear, their implementation can be complicated by the need for approximations. In this paper we review existing methods and then describe a fully Bayesian approach which allows approximation-free inference. We focus on predicting the failure time distribution of both future units and those that are currently under test. The methods are illustrated using fatigue crack growth data.

Bayes Theorem↗

Effectiveness of tibolone on the reduction of menopausal problems--a Bayesian semi-parametric interpretation.

Two major physiological problems women experience at the time of menopause are hot flush and vaginal dryness. Exploratory investigations reveal that these two binary outcomes are very much dependent as both of them have predominant oestrogenic effects. A primary interest is to investigate how the bivariate association and the marginal univariate risks are affected by repeated measurements on each woman over several months. To achieve this we propose a very general class of bivariate binary models. Parametric inference is drawn on the basis of full non-parametric Bayesian approach under Dirichlet process mixture. Study addresses some more interesting phenomena on the effectiveness of tibolone treatment in reducing menopausal problems. A simulation study further strengthens the proposed methodology.

Adult↗

Identification of co-regulated genes through Bayesian clustering of predicted regulatory binding sites.

The identification of co-regulated genes and their transcription-factor binding sites (TFBS) are key steps toward understanding transcription regulation. In addition to effective laboratory assays, various computational approaches for the detection of TFBS in promoter regions of coexpressed genes have been developed. The availability of complete genome sequences combined with the likelihood that transcription factors and their cognate sites are often conserved during evolution has led to the development of phylogenetic footprinting. The modus operandi of this technique is to search for conserved motifs upstream of orthologous genes from closely related species. The method can identify hundreds of TFBS without prior knowledge of co-regulation or coexpression. Because many of these predicted sites are likely to be bound by the same transcription factor, motifs with similar patterns can be put into clusters so as to infer the sets of co-regulated genes, that is, the regulons. This strategy utilizes only genome sequence information and is complementary to and confirmative of gene expression data generated by microarray experiments. However, the limited data available to characterize individual binding patterns, the variation in motif alignment, motif width, and base conservation, and the lack of knowledge of the number and sizes of regulons make this inference problem difficult. We have developed a Gibbs sampling-based Bayesian motif clustering (BMC) algorithm to address these challenges. Tests on simulated data sets show that BMC produces many fewer errors than hierarchical and K-means clustering methods. The application of BMC to hundreds of predicted gamma-proteobacterial motifs correctly identified many experimentally reported regulons, inferred the existence of previously unreported members of these regulons, and suggested novel regulons.

Algorithms↗

Bayesian restoration of ion channel records using hidden Markov models.

Hidden Markov models have been used to restore recorded signals of single ion channels buried in background noise. Parameter estimation and signal restoration are usually carried out through likelihood maximization by using variants of the Baum-Welch forward-backward procedures. This paper presents an alternative approach for dealing with this inferential task. The inferences are made by using a combination of the framework provided by Bayesian statistics and numerical methods based on Markov chain Monte Carlo stochastic simulation. The reliability of this approach is tested by using synthetic signals of known characteristics. The expectations of the model parameters estimated here are close to those calculated using the Baum-Welch algorithm, but the present methods also yield estimates of their errors. Comparisons of the results of the Bayesian Markov Chain Monte Carlo approach with those obtained by filtering and thresholding demonstrate clearly the superiority of the new methods.

Bayes Theorem↗

Randomization as a basis for inference in noninferiority trials.

Noninferiority testing in clinical trials is commonly understood in a Neyman-Pearson framework, and has been discussed in a Bayesian framework as well. In this paper, we discuss noninferiority testing in a Fisherian framework, in which the only assumption necessary for inference is the assumption of randomization of treatments to study subjects. Randomization plays an important role in not only the design but also the analysis of clinical trials, no matter the underlying inferential field. The ability to utilize permutation tests depends on assumptions around exchangeability, and we discuss the possible uses of permutation tests in active control noninferiority analyses. The other practical implications of this paper are admittedly minor but lead to better understanding of the historical and philosophical development of active control noninferiority testing. The conclusion may also frame discussion of other complicated issues in noninferiority testing, such as the role of an intention to treat analysis.

Clinical Trials as Topic↗

Codiversification in an ant-plant mutualism: stem texture and the evolution of host use in Crematogaster (Formicidae: Myrmicinae) inhabitants of Macaranga (Euphorbiaceae).

We investigate the evolution of host association in a cryptic complex of mutualistic Crematogaster (Decacrema) ants that inhabits and defends Macaranga trees in Southeast Asia. Previous phylogenetic studies based on limited samplings of Decacrema present conflicting reconstructions of the evolutionary history of the association, inferring both cospeciation and the predominance of host shifts. We use cytochrome oxidase I (COI) to reconstruct phylogenetic relationships in a comprehensive sampling of the Decacrema inhabitants of Macaranga. Using a published Macaranga phylogeny, we test whether the ants and plants have cospeciated. The COI phylogeny reveals 10 well-supported lineages and an absence of cospeciation. Host shifts, however, have been constrained by stem traits that are themselves correlated with Macaranga phylogeny. Earlier lineages of Decacrema exclusively inhabit waxy stems, a basal state in the Pachystemon clade within Macaranga, whereas younger species of Pachystemon, characterized by nonwaxy stems, are inhabited only by younger lineages of Decacrema. Despite the absence of cospeciation, the correlated succession of stem texture in both phylogenies suggests that Decacrema and Pachystemon have diversified in association, or codiversified. Subsequent to the colonization of the Pachystemon clade, Decacrema expanded onto a second clade within Macaranga, inducing the development of myrmecophytism in the Pruinosae group. Confinement to the aseasonal wet climate zone of western Malesia suggests myrmecophytic Macaranga are no older than the wet forest community in Southeast Asia, estimated to be about 20 million years old (early Miocene). Our calculation of COI divergence rates from several published arthropod studies that relied on tenable calibrations indicates a generally conserved rate of approximately 1.5% per million years. Applying this rate to a rate-smoothed Bayesian chronogram of the ants, the Decacrema from Macaranga are inferred to be at least 12 million years old (mid-Miocene). However, using the extremes of rate variation in COI produces an age as recent as 6 million years. Our inferred timeline based on 1.5% per million years concurs with independent biogeographical events in the region reconstructed from palynological data, thus suggesting that the evolutionary histories of Decacrema and their Pachystemon hosts have been contemporaneous since the mid-Miocene. The evolution of myrmecophytism enabled Macaranga to radiate into enemy-free space, while the ants' diversification has been shaped by stem traits, host specialization, and geographic factors. We discuss the possibility that the ancient and exclusive association between Decacrema and Macaranga was facilitated by an impoverished diversity of myrmecophytes and phytoecious (obligately plant inhabiting) ants in the region.

Animals↗