Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Hierarchical phylogenetic models for analyzing multipartite sequence data.

Debate exists over how to incorporate information from multipartite sequence data in phylogenetic analyses. Strict combined-data approaches argue for concatenation of all partitions and estimation of one evolutionary history, maximizing the explanatory power of the data. Consensus/independence approaches endorse a two-step procedure where partitions are analyzed independently and then a consensus is determined from the multiple results. Mixtures across the model space of a strict combined-data approach and a priori independent parameters are popular methods to integrate these methods. We propose an alternative middle ground by constructing a Bayesian hierarchical phylogenetic model. Our hierarchical framework enables researchers to pool information across data partitions to improve estimate precision in individual partitions while permitting estimation and testing of tendencies in across-partition quantities. Such across-partition quantities include the distribution from which individual topologies relating the sequences within a partition are drawn. We propose standard hierarchical priors on continuous evolutionary parameters across partitions, while the structure on topologies varies depending on the research problem. We illustrate our model with three examples. We first explore the evolutionary history of the guinea pig (Cavia porcellus) using alignments of 13 mitochondrial genes. The hierarchical model returns substantially more precise continuous parameter estimates than an independent parameter approach without losing the salient features of the data. Second, we analyze the frequency of horizontal gene transfer using 50 prokaryotic genes. We assume an unknown species-level topology and allow individual gene topologies to differ from this with a small estimable probability. Simultaneously inferring the species and individual gene topologies returns a transfer frequency of 17%. We also examine HIV sequences longitudinally sampled from HIV+ patients. We ask whether posttreatment development of CCR5 coreceptor virus represents concerted evolution from middisease CXCR4 virus or reemergence of initial infecting CCR5 virus. The hierarchical model pools partitions from multiple unrelated patients by assuming that the topology for each patient is drawn from a multinomial distribution with unknown probabilities. Preliminary results suggest evolution and not reemergence.

Animals↗

Structural inference in transition measurement error models for longitudinal data.

We propose a new class of models, transition measurement error models, to study the effects of covariates and the past responses on the current response in longitudinal studies when one of the covariates is measured with error. We show that the response variable conditional on the error-prone covariate follows a complex transition mixed effects model. The naive model obtained by ignoring the measurement error correctly specifies the transition part of the model, but misspecifies the covariate effect structure and ignores the random effects. We next study the asymptotic bias in naive estimator obtained by ignoring the measurement error for both continuous and discrete outcomes. We show that the naive estimator of the regression coefficient of the error-prone covariate is attenuated, while the naive estimators of the regression coefficients of the past responses are generally inflated. We then develop a structural modeling approach for parameter estimation using the maximum likelihood estimation method. In view of the multidimensional integration required by full maximum likelihood estimation, an EM algorithm is developed to calculate maximum likelihood estimators, in which Monte Carlo simulations are used to evaluate the conditional expectations in the E-step. We evaluate the performance of the proposed method through a simulation study and apply it to a longitudinal social support study for elderly women with heart disease. An additional simulation study shows that the Bayesian information criterion (BIC) performs well in choosing the correct transition orders of the models.

Aged↗

Phylogeny of the Procyonidae (Mammalia: Carnivora): molecules, morphology and the Great American Interchange.

The Procyonidae (Mammalia: Carnivora) have played a central role in resolving the controversial systematics of the giant and red pandas, but phylogenetic relationships of species within the family itself have received much less attention. Cladistic analyses of morphological characters conducted during the last two decades have resulted in topologies that group ecologically and morphologically similar taxa together. Specifically, the highly arboreal and frugivorous kinkajou (Potos flavus) and olingos (Bassaricyon) define one clade, whereas the more terrestrial and omnivorous coatis (Nasua), raccoons (Procyon), and ringtails (Bassariscus) define another clade, with the similar-sized Nasua and Procyon joined as sister taxa in this latter group. These relationships, however, have not been tested with molecular sequence data. We examined procyonid phylogenetics based on combined data from nine nuclear and two mitochondrial gene segments totaling 6534bp. We were able to fully resolve relationships within the family with strongly supported and congruent results from maximum parsimony, maximum likelihood, minimum evolution, and Bayesian analyses. We identified three distinct lineages within the family: a (Nasua, Bassaricyon) clade, a (Bassariscus, Procyon) clade, and a Potos lineage, the last of which is sister to the other two clades. These findings, which are in strong disagreement with prior fossil and morphology-based assessments of procyonid relationships, reemphasize the morphological and ecological flexibility of these taxa. In particular, morphological similarities between unrelated genera possibly reflect convergence associated with similar lifestyles and diets rather than ancestry. Furthermore, incongruence between the molecular supermatrix and a morphological character matrix comprised mostly of dental characters [Baskin, J.A., 2004. Bassariscus and Probassariscus (Mammalia, Carnivora, Procyonidae) from the early Barstovian (Middle Miocene). J. Vert. Paleo. 24, 709-720] may be due to non-independence among atomized dental characters that does not take into account the high developmental genetic correlation of these characters. Finally, molecular divergence dating analyses using a relaxed molecular clock approach suggest that intergeneric and intrageneric splits in the Procyonidae mostly occurred in the Miocene. The inferred divergence times for intrageneric splits for several genera whose ranges are bisected by the Panamanian Isthmus is significant because they suggest diversification well precedes the Great American Interchange, which has long been considered a primary underlying mechanism for procyonid evolution.

Animals↗

A test of geographic assignment using isotope tracers in feathers of known origin.

We used feathers of known origin collected from across the breeding range of a migratory shorebird to test the use of isotope tracers for assigning breeding origins. We analyzed deltaD, delta13C, and delta15N in feathers from 75 mountain plover (Charadrius montanus) chicks sampled in 2001 and from 119 chicks sampled in 2002. We estimated parameters for continuous-response inverse regression models and for discrete-response Bayesian probability models from data for each year independently. We evaluated model predictions with both the training data and by using the alternate year as an independent test dataset. Our results provide weak support for modeling latitude and isotope values as monotonic functions of one another, especially when data are pooled over known sources of variation such as sample year or location. We were unable to make even qualitative statements, such as north versus south, about the likely origin of birds using both deltaD and delta13C in inverse regression models; results were no better than random assignment. Probability models provided better results and a more natural framework for the problem. Correct assignment rates were highest when considering all three isotopes in the probability framework, but the use of even a single isotope was better than random assignment. The method appears relatively robust to temporal effects and is most sensitive to the isotope discrimination gradients over which samples are taken. We offer that the problem of using isotope tracers to infer geographic origin is best framed as one of assignment, rather than prediction.

Animal Migration↗

GenSo-FDSS: a neural-fuzzy decision support system for pediatric ALL cancer subtype identification using gene expression data.

OBJECTIVE: Acute lymphoblastic leukemia (ALL) is the most common malignancy of childhood, representing nearly one third of all pediatric cancers. Currently, the treatment of pediatric ALL is centered on tailoring the intensity of the therapy applied to a patient's risk of relapse, which is linked to the type of leukemia the patient has. Hence, accurate and correct diagnosis of the various leukemia subtypes becomes an important first step in the treatment process. Recently, gene expression profiling using DNA microarrays has been shown to be a viable and accurate diagnostic tool to identify the known prognostically important ALL subtypes. Thus, there is currently a huge interest in developing autonomous classification systems for cancer diagnosis using gene expression data. This is to achieve an unbiased analysis of the data and also partly to handle the large amount of genetic information extracted from the DNA microarrays. METHODOLOGY: Generally, existing medical decision support systems (DSS) for cancer classification and diagnosis are based on traditional statistical methods such as Bayesian decision theory and machine learning models such as neural networks (NN) and support vector machine (SVM). Though high accuracies have been reported for these systems, they fall short on certain critical areas. These included (a) being able to present the extracted knowledge and explain the computed solutions to the users; (b) having a logical deduction process that is similar and intuitive to the human reasoning process; and (c) flexible enough to incorporate new knowledge without running the risk of eroding old but valid information. On the other hand, a neural fuzzy system, which is synthesized to emulate the human ability to learn and reason in the presence of imprecise and incomplete information, has the ability to overcome the above-mentioned shortcomings. However, existing neural fuzzy systems have their own limitations when used in the design and implementation of DSS. Hence, this paper proposed the use of a novel neural fuzzy system: the generic self-organising fuzzy neural network (GenSoFNN) with truth-value restriction (TVR) fuzzy inference, as a fuzzy DSS (denoted as GenSo-FDSS) for the classification of ALL subtypes using gene expression data. RESULTS AND CONCLUSION: The performance of the GenSo-FDSS system is encouraging when benchmarked against those of NN, SVM and the K-nearest neighbor (K-NN) classifier. On average, a classification rate of above 90% has been achieved using the GenSo-FDSS system.

Algorithms↗

Geographical variations of inflammatory bowel disease in France: a study based on national health insurance data.

BACKGROUND AND AIM: A north-south gradient in inflammatory bowel disease (IBD) incidence has been found in Europe and the United States. Its existence is inferred from comparisons of registries that cover only small portions of territories. Several studies suggest that IBD incidence in the north has reached a plateau, whereas in the south it has risen sharply. This evolution tends to reduce the north-south gradient, and it is uncertain whether it still exists. In France, patients with IBD are fully reimbursed for their health expenses by the national health insurance system, which is a potential source of data concerning the incidence of IBD at the national level. The aim of this study was to assess the geographical distribution of Crohn's disease (CD) and ulcerative colitis (UC) in France and to test the north-south gradient hypothesis. METHODS: This study was conducted in metropolitan France and included patients to whom IBD reimbursement was newly attributed between January 1, 2000 and December 31, 2002. Data provided relate to age, sex, postcode area of residence, and IBD type. The mapping of geographical distribution of smoothed relative risks (RR) of CD and UC was carried out using a Bayesian approach, taking into account autocorrelation and population size in each département. RESULTS: In the overall population, incidence rates were 8.2 for CD and 7.2 for UC per 100,000 inhabitants. A clear north-south gradient was shown for CD. Départements with the highest smoothed RR were located in the northern third of France. By contrast, the geographical distribution of smoothed RR of UC was homogeneous. CONCLUSIONS: This study shows a north-south gradient in France for CD but not for UC.

Adolescent↗

Accounting for outliers and heteroskedasticity in multibreed genetic evaluations of postweaning gain of Nelore-Hereford cattle.

The objectives of this study were to demonstrate the utility of hierarchical Bayesian models combining residual heteroskedasticity with robustness for outlier detection and muting and to evaluate the effects of such joint modeling in multibreed genetic evaluations. A 3 x 2 factorial specification of 6 residual variance models based on several distributional (Gaussian, Student's t, or Slash) and variability (homoskedastic or heteroskedastic) assumptions was used to analyze 22,717 postweaning gain records from a Nelore-Hereford population (40,082 animals in the pedigree). To illustrate the utility of the 2 robust distributional specifications (Student's t and Slash) for outlier detection and muting, 3 records from the same contemporary group (an extreme residual outlier, a mild residual outlier, and a near-zero residual) were chosen for further study. The posterior densities of the corresponding weighting variables of these records were used to assess their degree of Gaussian outlyingness and the ability of the robust models to mute the effects of deviant records. The Student's t heteroskedastic provided the best-fit model among the 6 specifications and was preferred for genetic merit inference. Kendall rank correlations of the posterior means of the additive genetic effects of the animals, used to compare the selection order of the Student's t and Gaussian models, were reasonably high across all animals within the most frequent genotypes, ranging from 0.83 to 0.91 and from 0.89 to 0.95 for the homoskedastic and the heteroskedastic versions, respectively. However, when considering only animals ranked in the top 10% by the customary Gaussian homoskedastic model, these rank correlations were reduced considerably, ranging from 0.29 to 0.57 and from 0.72 to 0.85 between the 2 residual densities within the homoskedastic and heteroskedastic versions, respectively. Rank correlations between the homoskedastic and heteroskedastic versions within each of the Gaussian and Student's t error models tended to be smaller, with a range from 0.68 to 0.90 across all animals and from 0.28 to 0.67 for animals ranked in the top 10%. These results support the implementation of robust models accounting for sources of heteroskedasticity to increase the precision and stability of multibreed genetic evaluations with proper statistical treatment of deviant records.

Analysis of Variance↗

Fine-scale mapping of disease loci via shattered coalescent modeling of genealogies.

We present a Bayesian, Markov-chain Monte Carlo method for fine-scale linkage-disequilibrium gene mapping using high-density marker maps. The method explicitly models the genealogy underlying a sample of case chromosomes in the vicinity of a putative disease locus, in contrast with the assumption of a star-shaped tree made by many existing multipoint methods. Within this modeling framework, we can allow for missing marker information and for uncertainty about the true underlying genealogy and the makeup of ancestral marker haplotypes. A crucial advantage of our method is the incorporation of the shattered coalescent model for genealogies, allowing for multiple founding mutations at the disease locus and for sporadic cases of disease. Output from the method includes approximate posterior distributions of the location of the disease locus and population-marker haplotype proportions. In addition, output from the algorithm is used to construct a cladogram to represent genetic heterogeneity at the disease locus, highlighting clusters of case chromosomes sharing the same mutation. We present detailed simulations to provide evidence of improvements over existing methodology. Furthermore, inferences about the location of the disease locus are shown to remain robust to modeling assumptions.

Algorithms↗

Analysis of the overdispersed clock in the short-term evolution of hepatitis C virus: Using the E1/E2 gene sequences to infer infection dates in a single source outbreak.

The assumption of a molecular clock for dating events from sequence information is often frustrated by the presence of heterogeneity among evolutionary rates due, among other factors, to positively selected sites. In this work, our goal is to explore methods to estimate infection dates from sequence analysis. One such method, based on site stripping for clock detection, was proposed to unravel the clocklike molecular evolution in sequences showing high variability of evolutionary rates and in the presence of positive selection. Other alternatives imply accommodating heterogeneity in evolutionary rates at various levels, without eliminating any information from the data. Here we present the analysis of a data set of hepatitis C virus (HCV) sequences from 24 patients infected by a single individual with known dates of infection. We first used a simple criterion of relative substitution rate for site removal prior to a regression analysis. Time was regressed on maximum likelihood pairwise evolutionary distances between the sequences sampled from the source individual and infected patients. We show that it is indeed the fastest evolving sites that disturb the molecular clock and that these sites correspond to positively selected codons. The high computational efficiency of the regression analysis allowed us to compare the site-stripping scheme with random removal of sites. We demonstrate that removing the fast-evolving sites significantly increases the accuracy of estimation of infection times based on a single substitution rate. However, the time-of-infection estimations improved substantially when a more sophisticated and computationally demanding Bayesian method was used. This method was used with the same data set but keeping all the sequence positions in the analysis. Consequently, despite the distortion introduced by positive selection on evolutionary rates, it is possible to obtain quite accurate estimates of infection dates, a result of especial relevance for molecular epidemiology studies.

Bayes Theorem↗

Bayesian modeling of air pollution health effects with missing exposure data.

The authors propose a new statistical procedure that utilizes measurement error models to estimate missing exposure data in health effects assessment. The method detailed in this paper follows a Bayesian framework that allows estimation of various parameters of the model in the presence of missing covariates in an informative way. The authors apply this methodology to study the effect of household-level long-term air pollution exposures on lung function for subjects from the Southern California Children's Health Study pilot project, conducted in the year 2000. Specifically, they propose techniques to examine the long-term effects of nitrogen dioxide (NO2) exposure on children's lung function for persons living in 11 southern California communities. The effect of nitrogen dioxide exposure on various measures of lung function was examined, but, similar to many air pollution studies, no completely accurate measure of household-level long-term nitrogen dioxide exposure was available. Rather, community-level nitrogen dioxide was measured continuously over many years, but household-level nitrogen dioxide exposure was measured only during two 2-week periods, one period in the summer and one period in the winter. From these incomplete measures, long-term nitrogen dioxide exposure and its effect on health must be inferred. Results show that the method improves estimates when compared with standard frequentist approaches.

Air Pollutants↗

The footprints of visual attention during search with 100% valid and 100% invalid cues.

Human performance during visual search typically improves when spatial cues indicate the possible target locations. In many instances, the performance improvement is quantitatively predicted by a Bayesian or quasi-Bayesian observer in which visual attention simply selects the information at the cued locations without changing the quality of processing or sensitivity and ignores the information at the uncued locations. Aside from the general good agreement between the effect of the cue on model and human performance, there has been little independent confirmation that humans are effectively selecting the relevant information. In this study, we used the classification image technique to assess the effectiveness of spatial cues in the attentional selection of relevant locations and suppression of irrelevant locations indicated by spatial cues. Observers searched for a bright target among dimmer distractors that might appear (with 50% probability) in one of eight locations in visual white noise. The possible target location was indicated using a 100% valid box cue or seven 100% invalid box cues in which the only potential target locations was uncued. For both conditions, we found statistically significant perceptual templates shaped as differences of Gaussians at the relevant locations with no perceptual templates at the irrelevant locations. We did not find statistical significant differences between the shapes of the inferred perceptual templates for the 100% valid and 100% invalid cues conditions. The results confirm the idea that during search visual attention allows the observer to effectively select relevant information and ignore irrelevant information. The results for the 100% invalid cues condition suggests that the selection process is not drawn automatically to the cue but can be under the observers' voluntary control.

Attention↗

Phylogenetic structure of Zacco platypus (Teleostei, Cyprinidae) populations on the upper and middle Chang Jiang (=Yangtze) drainage inferred from cytochrome b sequences.

We examined the genetic structure and phylogenetic relationships of some Chinese populations from the Chang Jiang (=Yangtze) drainage of the cyprinid Zacco platypus. We sequenced the complete mitochondrial cytochrome b gene of 64 individuals from 6 upper and middle tributaries of the Sichuan and Hunan Provinces to assess their population structure and systematics. The combined analyses of the phylogenetic information and the population structure suggested that Chinese Z. platypus consist of four distinct mtDNA lineages which exhibit high genetic variation and haplotypic diversity (Zacco A-D). The high molecular divergence observed among Zacco A-D mtDNA lineages (TrN+I (0.76) distance, mean 8.9%+/-1.7%) and their phylogeographic structure indicate that all four lineages have evolved independently. Analysis of molecular variance (AMOVA) indicates that most of the genetic variation observed is found among the four Zacco mtDNA lineages (thetaCT = 0.94) suggesting restricted gene flow among the Chang Jiang populations. Long-term interruption of gene flow was also evidenced by thetaST values higher than 0.9 that could be favoured by the discontinuous distributions of the lineages inhabiting upper (Sichuan Province) and middle (Hunan Province) Chang Jiang tributaries. The significant correlation between the geographic and genetic distances provide support for the importance of geographic discontinuity in shaping the Zacco genetic structure. Nested clade analysis (NCA) results were congruent with phylogenetic relationships recovered and confirm the genetic distinctiveness of four independent Zacco groups. These four groups correspond to the four Zacco A-D mtDNA lineages recovered in the phylogeny and were defined by nucleotide synapomorphies permitting bootstrapped and Bayesian confidence of 95% or greater. The high level of mitochondrial sequence divergence separating all Zacco mtDNA lineages suggested that the Z. platypus populations from the Chang Jiang drainage probably correspond to four different species.

Animals↗

Tree and rate estimation by local evaluation of heterochronous nucleotide data.

MOTIVATION: Heterochronous gene sequence data is important for characterizing the evolutionary processes of fast-evolving organisms such as RNA viruses. A limited set of algorithms exists for estimating the rate of nucleotide substitution and inferring phylogenetic trees from such data. The authors here present a new method, Tree and Rate Estimation by Local Evaluation (TREBLE) that robustly calculates the rate of nucleotide substitution and phylogeny with several orders of magnitude improvement in computational time. METHODS: For the basis of its rate estimation TREBLE novelly utilizes a geometric interpretation of the molecular clock assumption to deduce a local estimate of the rate of nucleotide substitution for triplets of dated sequences. Averaging the triplet estimates via a variance weighting yields a global estimate of the rate. From this value, an iterative refinement procedure relying on statistical properties of the triplets then generates a final estimate of the global rate of nucleotide substitution. The estimated global rate is then utilized to find the tree from the pairwise distance matrix via an UPGMA-like algorithm. RESULTS: Simulation studies show that TREBLE estimates the rate of nucleotide substitution with point estimates comparable with the best of available methods. Confidence intervals are comparable with that of BEAST. TREBLE's phylogenetic reconstruction is significantly improved over the other distance matrix method but not as accurate as the Bayesian algorithm. Compared with three other algorithms, TREBLE reduces computational time by a minimum factor of 3000. Relative to the algorithm with the most accurate estimates for the rate of nucleotide substitution (i.e. BEAST), TREBLE is over 10,000 times more computationally efficient. AVAILABILITY: jdobrien.bol.ucla.edu/TREBLE.html

Chromosome Mapping↗

Bootstrap, Bayesian probability and maximum likelihood mapping: exploring new tools for comparative genome analyses.

BACKGROUND: Horizontal gene transfer (HGT) played an important role in shaping microbial genomes. In addition to genes under sporadic selection, HGT also affects housekeeping genes and those involved in information processing, even ribosomal RNA encoding genes. Here we describe tools that provide an assessment and graphic illustration of the mosaic nature of microbial genomes. RESULTS: We adapted the Maximum Likelihood (ML) mapping to the analyses of all detected quartets of orthologous genes found in four genomes. We have automated the assembly and analyses of these quartets of orthologs given the selection of four genomes. We compared the ML-mapping approach to more rigorous Bayesian probability and Bootstrap mapping techniques. The latter two approaches appear to be more conservative than the ML-mapping approach, but qualitatively all three approaches give equivalent results. All three tools were tested on mitochondrial genomes, which presumably were inherited as a single linkage group. CONCLUSIONS: In some instances of interphylum relationships we find nearly equal numbers of quartets strongly supporting the three possible topologies. In contrast, our analyses of genome quartets containing the cyanobacterium Synechocystis sp. indicate that a large part of the cyanobacterial genome is related to that of low GC Gram positives. Other groups that had been suggested as sister groups to the cyanobacteria contain many fewer genes that group with the Synechocystis orthologs. Interdomain comparisons of genome quartets containing the archaeon Halobacterium sp. revealed that Halobacterium sp. shares more genes with Bacteria that live in the same environment than with Bacteria that are more closely related based on rRNA phylogeny. Many of these genes encode proteins involved in substrate transport and metabolism and in information storage and processing. The performed analyses demonstrate that relationships among prokaryotes cannot be accurately depicted by or inferred from the tree-like evolution of a core of rarely transferred genes; rather prokaryotic genomes are mosaics in which different parts have different evolutionary histories. Probability mapping is a valuable tool to explore the mosaic nature of genomes.

Journal Article↗

Accounting for uncertainty in the tree topology has little effect on the decision-theoretic approach to model selection in phylogeny estimation.

Currently available methods for model selection used in phylogenetic analysis are based on an initial fixed-tree topology. Once a model is picked based on this topology, a rigorous search of the tree space is run under that model to find the maximum-likelihood estimate of the tree (topology and branch lengths) and the maximum-likelihood estimates of the model parameters. In this paper, we propose two extensions to the decision-theoretic (DT) approach that relax the fixed-topology restriction. We also relax the fixed-topology restriction for the Bayesian information criterion (BIC) and the Akaike information criterion (AIC) methods. We compare the performance of the different methods (the relaxed, restricted, and the likelihood-ratio test [LRT]) using simulated data. This comparison is done by evaluating the relative complexity of the models resulting from each method and by comparing the performance of the chosen models in estimating the true tree. We also compare the methods relative to one another by measuring the closeness of the estimated trees corresponding to the different chosen models under these methods. We show that varying the topology does not have a major impact on model choice. We also show that the outcome of the two proposed extensions is identical and is comparable to that of the BIC, Extended-BIC, and DT. Hence, using the simpler methods in choosing a model for analyzing the data is more computationally feasible, with results comparable to the more computationally intensive methods. Another outcome of this study is that earlier conclusions about the DT approach are reinforced. That is, LRT, Extended-AIC, and AIC result in more complicated models that do not contribute to the performance of the phylogenetic inference, yet cause a significant increase in the time required for data analysis.

Computational Biology↗

Bayesian analysis via Gibbs sampling of susceptibility to intramammary infection in Holstein cattle.

A Bayesian analysis was undertaken to assess the susceptibility of Holsteins to mastitis from 120 to 305 d in milk. Data included 595 lactations from 267 cows. The response variable was presence or absence of intramammary infection; explanatory variables were period and season of calving, somatic cell score, and cow. The logistic model adopted had period and season of calving and the regression on somatic cell score with vague prior distributions, and cow effects had a normal prior with unknown variance sigma u2, which, in turn, had a gamma prior. Implementation was by Gibbs sampling. Posterior densities of location parameters were unimodal and symmetric. The probability of intramammary infection of a sample cow was skewed. The posterior distribution of sigma u2 was skewed also. Gibbs samples of sigma u2 had high lag correlations, which gave an effective sample ranging between 47 and 117 from a chain of size 3000. There were differences between estimates of sigma u2 found using Gibbs sampling and those obtained using approximations. The low information content arising from the small size of the data and the binary nature of the response are reasons for such differences. A sensitivity analysis revealed influences of hyperparameters of the prior distribution of sigma u2 on inferences about this parameter.

Animals↗

Statistical issues in the analysis of disease mapping data.

In this paper we discuss a number of issues that are pertinent to the analysis of disease mapping data. As an illustrative example we consider the mapping of larynx cancer across electoral wards in the North West Thames region of the U.K. Bayesian hierarchical models are now frequently employed to carry out such mapping. In a typical situation, a three-stage hierarchical model is specified in which the data are modelled as a function of area-specific relative risks at stage one; the collection of relative risks across the study region are modelled at stage two; and at stage three prior distributions are assigned to parameters of the stage two distribution. Such models allow area-specific disease relative risks to be 'smoothed' towards global and/or local mean levels across the study region. However, these models contain many structural and functional assumptions at different levels of the hierarchy; we aim to discuss some of these assumptions and illustrate their sensitivity. When relative risks are the endpoint of interest, it is common practice to assume that, for each of the age-sex strata of a particular area, there is a common multiplier (the relative risk) acting upon each of the stratum-specific risks in that area; we will examine this proportionality assumption. We also consider the choices of models and priors at stages two and three of the hierarchy, the effect of outlying areas, and an assessment of the level of smoothing that is being carried out. For inference, we concentrate on the description of the spatial variability in relative risks and on the association between the relative risks of larynx cancer and an area-level measure of socio-economic status.

Age Factors↗

Bayesian non-response models for categorical data from small areas: an application to BMD and age.

We provide a Bayesian analysis of data categorized into two levels of age (younger than 50 years, at least 50 years) and three levels of bone mineral density (normal, osteopenia, osteoporosis) for white females at least 20 years old in the third National Health and Nutrition Examination Survey. For the sample, the age of each individual is known, but some individuals did not have their BMD measured. We use two types of models: In the ignorable non-response models the propensity to respond does not depend on BMD and age of an individual, while in the non-ignorable non-response models it does. These are the baseline models which are used to derive all models for testing. Our non-ignorable non-response models are 'close' to the ignorable non-response models, thereby reducing the effects of the assumptions about non-respondents that cannot be tested in non-response models. We have data from 35 counties, small areas, and therefore our models are hierarchical, a feature that allows a 'borrowing of strength' across the counties, and they provide a substantial reduction in variation. The non-ignorable non-response models are generalizations of the ignorable non-response models, and therefore, the non-ignorable non-response models allow broader inference. The joint posterior density of the parameters for each model is complex, and therefore, we fit each model using Markov chain Monte Carlo methods to obtain samples which are used to make inference about BMD and age. For each county we can estimate the proportion of individuals in each BMD and age cell of the categorical table, and we can assess the relation between BMD and age using the Bayes factor. A sensitivity analysis shows that there are differences (typically small) in inference that permits different levels of association between BMD and age. A simulation study shows that there is not much difference between the baseline ignorable and non-ignorable non-response models.

Adult↗