Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Lidar detection of underwater objects using a neuro-SVM-based architecture.

This paper presents a neural network architecture using a support vector machine (SVM) as an inference engine (IE) for classification of light detection and ranging (Lidar) data. Lidar data gives a sequence of laser backscatter intensities obtained from laser shots generated from an airborne object at various altitudes above the earth surface. Lidar data is pre-filtered to remove high frequency noise. As the Lidar shots are taken from above the earth surface, it has some air backscatter information, which is of no importance for detecting underwater objects. Because of these, the air backscatter information is eliminated from the data and a segment of this data is subsequently selected to extract features for classification. This is then encoded using linear predictive coding (LPC) and polynomial approximation. The coefficients thus generated are used as inputs to the two branches of a parallel neural architecture. The decisions obtained from the two branches are vector multiplied and the result is fed to an SVM-based IE that presents the final inference. Two parallel neural architectures using multilayer perception (MLP) and hybrid radial basis function (HRBF) are considered in this paper. The proposed structure fits the Lidar data classification task well due to the inherent classification efficiency of neural networks and accurate decision-making capability of SVM. A Bayesian classifier and a quadratic classifier were considered for the Lidar data classification task but they failed to offer high prediction accuracy. Furthermore, a single-layered artificial neural network (ANN) classifier was also considered and it failed to offer good accuracy. The parallel ANN architecture proposed in this paper offers high prediction accuracy (98.9%) and is found to be the most suitable architecture for the proposed task of Lidar data classification.

Algorithms↗

A Bayesian hierarchical survival model for the institutional effects in a multi-centre cancer clinical trial.

In randomized clinical trials comparing treatment effects on diseases such as cancer, a multi-centre trial is usually conducted to accrue the required number of patients in a reasonable period of time. While we interpret the average treatment effect, it is necessary to examine the homogeneity of the observed treatment effects across institutions, that is, treatment-by-institution interaction. If the homogeneity is confirmed, the conclusions concerning treatment effects can be generalized to a broader patient population. In this paper, a Bayesian hierarchical survival model is used to investigate the institutional effects on the efficacy of treatment as well as on the baseline risk. The marginal posterior distributions are estimated by a Markov chain Monte Carlo method, that is, Gibbs sampling, to overcome current computational limitations. The robustness of the inferences to the distributional assumption for the random effects is also examined. We illustrate the methods with analyses of data from a multi-centre cancer clinical trial, which investigated the efficacy of immunochemotherapy as an adjuvant treatment after curative resection of gastric cancer. In this trial there is little difference in the treatment effects across institutions and the treatment is shown to be effective, while there appears to be substantial variation in the baseline risk across institutions. This result indicates that the observed treatment effects might be generalized to a broader patient population.

Adjuvants, Immunologic↗

Statistical assessment of mediational effects for logistic mediational models.

The concept of mediation has broad applications in medical health studies. Although the statistical assessment of a mediational effect under the normal assumption has been well established in linear structural equation models (SEM), it has not been extended to the general case where normality is not a usual assumption. In this paper, we propose to extend the definition of mediational effects through causal inference. The new definition is consistent with that in linear SEM and does not rely on the assumption of normality. Here, we focus our attention on the logistic mediation model, where all variables involved are binary. Three approaches to the estimation of mediational effects-Delta method, bootstrap, and Bayesian modelling via Monte Carlo simulation are investigated. Simulation studies are used to examine the behaviour of the three approaches. Measured by 95 per cent confidence interval (CI) coverage rate and root mean square error (RMSE) criteria, it was found that the Bayesian method using a non-informative prior outperformed both bootstrap and the Delta methods, particularly for small sample sizes. Case studies are presented to demonstrate the application of the proposed method to public health research using a nationally representative database. Extending the proposed method to other types of mediational model and to multiple mediators are also discussed.

Adolescent↗

Dengue virus circulation and evolution in Mexico: a phylogenetic perspective.

BACKGROUND: Dengue is the most important arthropod-borne viral infection in the Americas. In the last decades a progressive increment in dengue severity has been observed in Mexico and other countries of the region. METHODS: Molecular epidemiological studies were conducted to investigate the viral determinants of the emergence of epidemic dengue, dengue hemorrhagic fever and dengue shock syndrome as major public health problems in Mexico. Bayesian phylogenetic analyses were conducted to determine the origin, persistence and geographical dispersion of the four serotypes of dengue virus (DENV) isolated in Mexico between 1980 and 2002. Tests for natural selection were also conducted. RESULTS: The origin of some, but not all, strains circulating in Mexico could be inferred. Frequent lineage replacements were observed and were likely due to stochastic events. In situ evolution was detected but not associated with natural selection. Recent changes in the incidence and severity of dengue were temporally associated with the introduction and circulation of different serotypes and genotypes of DENV. CONCLUSIONS: Introduction of new DENV genotypes and serotypes is a major risk factor for epidemic dengue and severe disease. Increased surveillance for such introductions is critical to allow public health authorities to intervene in impending epidemics.

Aedes↗

A Bayesian framework for multilead SMD post-placement quality inspection.

In this paper, a novel framework is proposed to inspect the placement quality of surface mount technology devices (SMDs), immediately after they have been placed in wet solder paste on a printed circuit board (PCB). The developed approach involves the indirect measurement of each lead displacement with respect to its ideal position, centralized on its pad region. This displacement is inferred from area measurements on the raw image data of the lead region through a classification process. To increase the accuracy in the computation of the lead displacement, we introduce a combined classification/estimation process, in which the individual lead displacement classifications are viewed as measurements (or observations) of the same physical quantity i.e., the displacement of the entire component as a rigid body. Certain geometric relations connecting lead shifts to component displacement are also derived. Employing these relations we can infer a new refined measurement of the shift of each individual lead, a quantity crucial to the calculation of the quality measures. Experimental results highlight the potential of the developed algorithm.

Journal Article↗

On the analysis of accumulation curves.

Identifying and counting the total number of biological species observed, when plotted against a measure of the effort used to record them, gives rise to a species accumulation curve. We investigate estimation of the total number of species and other relevant properties of accumulation by elaborating on the multinomial model (Nakamura and Peraza, 1998, Journal of Agricultural, Biological, and Environmental Statistics, 3, 17-36) that includes specification of a beta density for recording probabilities. We consider a unified description including complete and incomplete (aggregated) curves, a more general scheme for recording, and a Bayesian framework to allow for inclusion of the biologist's knowledge with regard to typical recording probabilities and the total number of species. The beta distribution is used as a prior, but recording probabilities are not restrained to be beta distributed. The methods yield either closed analytical expressions for inference or relatively simple numerical procedures. Predictive distributions of future recordings that would eventually lead to an optimal decision-theoretical rule for stopping collection effort may be easily calculated. A case study regarding species of bats is considered, including some guidelines for elicitation of an informative prior.

Animals↗

Projection of lung cancer mortality in Japan.

According to the National Vital Statistics data, age-standardized mortality rates (ASRs) of lung cancer have shown slightly declining trends in Japan for both men and women. In order to evaluate whether this tendency will continue, a Bayesian age-period-cohort (APC) model was applied using the National Vital Statistics data from 1952 to 2001. In the projection, a Gaussian autoregressive prior model was applied to smooth age, period, and cohort effects from its 2 immediate predecessors by extrapolation. Posterior distributions from which we drew inferences on mortality rates were derived from 15,000 iterations using 5000 burn-in iterations. We defined the median of the iterated values as the overall summary mortality rate of the iterated results. Our results suggest that the number of deaths due to lung cancer will double for men and women during the next 3 decades due to the aging of the baby-boomer generation (individuals who were born between 1947 and 1951). Currently declining trends in some age groups will reverse and start to increase again in the next decades. However, for recent birth cohorts, the results of the projection varied according to whether the data set included early age group mortality or not. Lung cancer mortality in the future depends on the risk factors engaged in by today's young people, especially smoking. Strong promotion of anti-smoking measures and careful surveillance for lung cancer are needed.

Adult↗

Genuine Bayesian multiallelic significance test for the Hardy-Weinberg equilibrium law.

Statistical tests that detect and measure deviation from the Hardy-Weinberg equilibrium (HWE) have been devised but are limited when testing for deviation at multiallelic DNA loci is attempted. Here we present the full Bayesian significance test (FBST) for the HWE. This test depends neither on asymptotic results nor on the number of possible alleles for the particular locus being evaluated. The FBST is based on the computation of an evidence index in favor of the HWE hypothesis. A great deal of forensic inference based on DNA evidence assumes that the HWE is valid for the genetic loci being used. We applied the FBST to genotypes obtained at several multiallelic short tandem repeat loci during routine parentage testing; the locus Penta E exemplifies those clearly in HWE while others such as D10S1214 and D19S253 do not appear to show this.

Alleles↗

Phylogenetic relationships and convergence of helicosporous fungi inferred from ribosomal DNA sequences.

Helicosporous fungi form elegant, coiled, and multicellular mitotic spores (conidia). In this paper, we investigate the phylogenetic relationships among helicosporous fungi in the asexual genera Helicoma, Helicomyces, Helicosporium, Helicodendron, Helicoon, and in the sexual genus Tubeufia (Tubeufiaceae, Dothideomycetes, and Ascomycota). We generated ribosomal small subunit and partial large subunit sequences from 39 fungal cultures. These and related sequences from GenBank were analyzed using parsimony, likelihood, and Bayesian analysis. Results showed that helicosporous species arose convergently from six lineages of fungi in the Ascomycota. The Tubeufiaceae s. str. formed a strongly supported monophyletic lineage comprising most species from Helicoma, Helicomyces, and Helicosporium. However, within the Tubeufiaceae, none of the asexual genera were monophyletic. Traditional generic characters, such as whether conidiophores were conspicuous or reduced, the thickness of the conidial filament, and whether or not conidia were hygroscopic, were more useful for species delimitation than for predicting higher level relationships. In spite of their distinctive, barrel-shaped spores, Helicoon species were polyphyletic and had evolved in different ascomycete orders. Helicodendron appeared to be polyphyletic although most representatives occurred within Leotiomycetes. We speculate that some of the convergent spore forms may represent adaptation to dispersal in aquatic environments.

DNA, Ribosomal↗

Phylogenetic relationships of Steinernema Travassos, 1927 (Nematoda: Cephalobina: Steinernematidae) based on nuclear, mitochondrial and morphological data.

Entomopathogenic nematodes of the genus Steinernema are lethal parasites of insects that are used as biological control agents of several lepidopteran, dipteran and coleopteran pests. Phylogenetic relationships among 25 Steinernema species were estimated using nucleotide sequences from three genes and 22 morphological characters. Parsimony analysis of 28S (LSU) sequences yielded a well-resolved phylogenetic hypothesis with reliable bootstrap support for 13 clades. Parsimony analysis of mitochondrial DNA sequences (12S rDNA and cox 1 genes) yielded phylogenetic trees with a lower consistency index than for LSU sequences, and with fewer reliably supported clades. Combined phylogenetic analysis of the 3-gene dataset by parsimony and Bayesian methods yielded well-resolved and highly similar trees. Bayesian posterior probabilities were high for most clades; bootstrap (parsimony) support was reliable for approximately half of the internal nodes. Parsimony analysis of the morphological dataset yielded a poorly resolved tree, whereas total evidence analysis (molecular plus morphological data) yielded a phylogenetic hypothesis consistent with, but less resolved than trees inferred from combined molecular data. Parsimony mapping of morphological characters on the 3-gene trees showed that most structural features of steinernematids are highly homoplastic. The distribution of nematode foraging strategies on these trees predicts that S. hermaphroditum, S. diaprepesi and S. longicaudum (US isolate) have cruise forager behaviours.

Amino Acid Sequence↗

Molecular phylogeny of penaeid shrimps inferred from two mitochondrial markers.

Penaeid shrimps are an important resource in crustacean fisheries, representing more than the half of the gross production of shrimp worldwide. In the present study, we used a sample of wide-ranging diversity (41 shrimp species) and two mitochondrial markers (758 bp) to clarify the evolutionary relationships among Penaeidae genera. Three different methodologies of tree reconstruction were employed in the study: maximum likelihood, neighbor joining and Bayesian analysis. Our results suggest that the old Penaeus genus is monophyletic and that the inclusion of the Solenocera genus within the Penaeidae family remains uncertain. With respect to Metapenaeopsis monophyly, species of this genus appeared clustered, but with a nonsignificant bootstrap value. These results elucidate some features of the unclear evolution of Penaeidae and may contribute to the taxonomic characterization of this family.

Algorithms↗

A molecular phylogeny of the Canidae based on six nuclear loci.

We have reconstructed the phylogenetic relationships of 23 species in the dog family, Canidae, using DNA sequence data from six nuclear loci. Individual gene trees were generated with maximum parsimony (MP) and maximum likelihood (ML) analysis. In general, these individual gene trees were not well resolved, but several identical groupings were supported by more than one locus. Phylogenetic analysis with a data set combining the six nuclear loci using MP, ML, and Bayesian approaches produced a more resolved tree that agreed with previously published mitochondrial trees in finding three well-defined clades, including the red fox-like canids, the South American foxes, and the wolf-like canids. In addition, the nuclear data set provides novel indel support for several previously inferred clades. Differences between trees derived from the nuclear data and those from the mitochondrial data include the grouping of the bush dog and maned wolf into a clade with the South American foxes, the grouping of the side-striped jackal (Canis adustus) and black-backed jackal (Canis mesomelas) and the grouping of the bat-eared fox (Otocyon megalotis) with the raccoon dog (Nycteruetes procyonoides). We also analyzed the combined nuclear+mitochondrial tree. Many nodes that were strongly supported in the nuclear tree or the mitochondrial tree remained strongly supported in the nuclear+mitochondrial tree. Relationships within the clades containing the red fox-like canids and South American canids are well resolved, whereas the relationships among the wolf-like canids remain largely undetermined. The lack of resolution within the wolf-like canids may be due to their recent divergence and insufficient time for the accumulation of phylogenetically informative signal.

Animals↗

A case study on the choice, interpretation and checking of multilevel models for longitudinal binary outcomes.

Recent advances in statistical software have led to the rapid diffusion of new methods for modelling longitudinal data. Multilevel (also known as hierarchical or random effects) models for binary outcomes have generally been based on a logistic-normal specification, by analogy with earlier work for normally distributed data. The appropriate application and interpretation of these models remains somewhat unclear, especially when compared with the computationally more straightforward semiparametric or 'marginal' modelling (GEE) approaches. In this paper we pose two interrelated questions. First, what limits should be placed on the interpretation of the coefficients and inferences derived from random-effect models involving binary outcomes? Second, what diagnostic checks are appropriate for evaluating whether such random-effect models provide adequate fits to the data? We address these questions by means of an extended case study using data on adolescent smoking from a large cohort study. Bayesian estimation methods are used to fit a discrete-mixture alternative to the standard logistic-normal model, and posterior predictive checking is used to assess model fit. Surprising parallels in the parameter estimates from the logistic-normal and mixture models are described and used to question the interpretability of the so-called 'subject-specific' regression coefficients from the standard multilevel approach. Posterior predictive checks suggest a serious lack of fit of both multilevel models. The results do not provide final answers to the two questions posed, but we expect that lessons learned from the case study will provide general guidance for further investigation of these important issues.

Journal Article↗

A statistical framework for quantitative trait mapping.

We describe a general statistical framework for the genetic analysis of quantitative trait data in inbred line crosses. Our main result is based on the observation that, by conditioning on the unobserved QTL genotypes, the problem can be split into two statistically independent and manageable parts. The first part involves only the relationship between the QTL and the phenotype. The second part involves only the location of the QTL in the genome. We developed a simple Monte Carlo algorithm to implement Bayesian QTL analysis. This algorithm simulates multiple versions of complete genotype information on a genomewide grid of locations using information in the marker genotype data. Weights are assigned to the simulated genotypes to capture information in the phenotype data. The weighted complete genotypes are used to approximate quantities needed for statistical inference of QTL locations and effect sizes. One advantage of this approach is that only the weights are recomputed as the analyst considers different candidate models. This device allows the analyst to focus on modeling and model comparisons. The proposed framework can accommodate multiple interacting QTL, nonnormal and multivariate phenotypes, covariates, missing genotype data, and genotyping errors in any type of inbred line cross. A software tool implementing this procedure is available. We demonstrate our approach to QTL analysis using data from a mouse backcross population that is segregating multiple interacting QTL associated with salt-induced hypertension.

Algorithms↗

Improving phylogenetic inference of mushrooms with RPB1 and RPB2 nucleotide sequences (Inocybe; Agaricales).

Approximately 3000 bp across 84 taxa have been analyzed for variable regions of RPB1, RPB2, and nLSU-rDNA to infer phylogenetic relationships in the large ectomycorrhizal mushroom genus Inocybe (Agaricales; Basidiomycota). This study represents the first effort to combine variable regions of RPB1 and RPB2 with nLSU-rDNA for low-level phylogenetic studies in mushroom-forming fungi. Combination of the three loci increases non-parametric bootstrap support, Bayesian posterior probabilities, and resolution for numerous clades compared to separate gene analyses. These data suggest the evolution of at least five major lineages in Inocybe-the Inocybe clade, the Mallocybe clade, the Auritella clade, the Inosperma clade, and the Pseudosperma clade. Additionally, many clades nested within each major lineage are strongly supported. These results also suggest the family Crepiodataceae sensu stricto is sister to Inocybe. Recognition of Inocybe at the family level, the Inocybaceae, is recommended.

Agaricales↗

Dual screening.

We discuss the problem of screening a general population for characteristics such as HIV or drug use. Our main approach is Bayesian, which allows for the incorporation of prior information about parameters. In the particular problem we consider, there is currently no information in the data for estimating the sensitivity of the screening test, and consequently, the prevalence of the characteristic among screened negatives cannot be estimated from the collected data alone. Our inferences are straightforward to obtain using Gibbs sampling techniques, and they are valid for large or small samples and for arbitrary prevalence or accuracy of screening tests. We also develop the maximum-likelihood approach using the EM algorithm.

AIDS Serodiagnosis↗

Genetic relatedness and population differentiation of Himalayan hulless barley (Hordeum vulgare L.) landraces inferred with SSRs.

A set of 107 hulless barley (Hordeum vulgare L. subsp. vulgare) landraces originally collected from the highlands of Nepal along the Annapurna and Manaslu Himalaya range were studied for genetic relatedness and population differentiation using simple sequence repeats (SSRs). The 44 genome covering barley SSRs applied in this study revealed a high level of genetic diversity among the landraces (diversity index, DI = 0.536) tested. The genetic similarity (GS) based UPGMA clustering and Bayesian Model-based (MB) structure analysis revealed a complex genetic structure of the landraces. Eight genetically distinct populations were identified, of which seven were further studied for diversity and differentiation. The genetic diversity estimated for all and each population separately revealed a hot spot of genetic diversity at Pisang (DI = 0.559). The populations are fairly differentiated (theta = 0.433, R(ST) = 0.445) accounting for > 40% of the genetic variation among the populations. The pairwise population differentiation test confirmed that many of the geographic populations significantly differ from each other but that the differentiation is independent of the geographic distance (r = 0.224, P > 0.05). The high level of genetic diversity and complex population structure detected in Himalayan hulless barley landraces and the relevance of the findings are discussed.

Alleles↗

On the logic of hypothesis testing in functional imaging.

Statistics is nowadays the customary language of functional imaging. It is common to express an experimental setting as a set of null hypotheses over complex models and to present results as maps of p-values derived from sophisticated probability distributions. However, the growing interest in the development of advanced statistical algorithms is not always paralleled by similar attention to how these techniques may regiment the ways in which users draw inferences from their data. This article investigates the logical bases of current statistical approaches in functional imaging and probes their suitability to inductive inference in neuroscience. The frequentist approach to statistical inference is reviewed with attention to its two main constituents: Fisherian "significance testing" and Neyman-Pearson "hypothesis testing". It is shown that these conceptual systems, which are similar in the univariate testing case, dissociate into two quite different methods of inference when applied to the multiple testing problem, the typical framework of functional imaging. This difference is explained with reference to specific issues, like small volume correction, which are most likely to generate confusion in the practitioner. Further insight into this problem is achieved by recasting the multiple comparison problem into a multivariate Bayesian formulation. This formulation introduces a new perspective where the inferential process is more clearly defined in two distinct steps. The first one, inductive in form, uses exploratory techniques to acquire preliminary notions on the spatial patterns and the signal and noise characteristics. The (smaller) set of likely spatial patterns generated is then tested with newer data and a more rigorous multiple hypothesis testing technique (deductive step).

Algorithms↗