The computer package DismapWin.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
In studying geographic disease distributions, one normally compares rates among arbitrarily defined geographic subareas (for example, census tracts), thereby sacrificing the geographic detail of the original data. The sparser the data, the larger the subareas must be in order to calculate stable rates. This dilemma is avoided with the technique of density equalizing map projections (DEMP). Boundaries of geographic subregions are adjusted to equalize population density over the entire study area. Case location plotted on the transformed map should have a uniform distribution if the underlying disease rates are constant. The present report describes the application of the DEMP technique to 401 childhood cancer cases occurring between 1980 and 1988 in four California counties, with the use of map files and population data for the 262 tracts of the 1980 Census. A kth nearest neighbour analysis provides strong evidence for geographic non-uniformity in tract rates (p < 10(-4)). No such effect is observed for artificial cases generated under the assumption of constant rates. Work is in progress to repeat the analysis with improved population estimates derived from both 1980 and 1990 Census data. Final epidemiologic conclusions will be reported when that analysis is complete.
Advances in computer hardware, software and database interfaces have provided opportunities for collation, manipulation, analysis and display of spatial data on an unprecedented scale. Demands for small area data in public health, fuelled in part by an increasing emphasis on benchmarking in relation to Year 2000 objectives but also in response to state and federal programmes to involve local communities in the assessment and planning process have simultaneously generated an unprecedented demand for these analyses and data presentations. This paper discusses four areas where geographic, cartographic and statistical theory and methodology need to be brought to bear on the development of applications involving small area health data. These areas are: (i) the theoretical conceptions of space; (ii) managing the inherent variability of rates and frequencies; (iii) attribution of events or cases to areas or to points; and (iv) the application of sound principles of cartographic design to the presentation of results.
Geographic information systems (GIS) and digital computer technology will advance the mission of the Centers for Disease Control and Prevention (CDC) and Agency for Toxic Substances and Disease Registry (ATSDR) to protect public health. Geographic positioning, topology, and planar and surface measurements are basic GIS properties which enable highly precise locational referencing of spatial phenomena. The growing uses of remotely sensed imagery and satellite facilitated global positioning systems are contributing to unprecedented surveillance of the environment and greater understanding of known and suspected environmental disease associations with human and animal health. Earth science and public health monitoring GIS databases offer new analytic opportunities for disease assessment and prevention.
We describe Bayesian hierarchical-spatial models for disease mapping with imprecisely observed ecological covariates. We posit smoothing priors for both the disease submodel and the covariate submodel. We apply the models to an analysis of insulin Dependent Diabetes Mellitus incidence in Sardinia, with malaria prevalence as a covariate.
This paper considers the underlying principles of depicting disease incidence on geographical maps and uses them to attempt a comparative classification of methods. After a discussion of the possibilities for incorporating time, we consider projection methods, some of which have been used to portray information in a manner supposed to be independent of population density. We then distinguish between non-parametric and model-based methods, including models for areal data using Bayesian ideas. Data in point form are also discussed and it is argued that the relative risk function provides a fundamental model useful for assessing different methods as a whole, some of which are known to be flawed and many of which are untested as regards their statistical properties.
The analysis of small area disease incidence has now developed to a degree where many methods have been proposed. However, there are few studies of the relative merits of the methods available. While many Bayesian models have been examined with respect to prior sensitivity, it is clear that wider comparisons of methods are largely missing from the literature. In this paper we present some preliminary results concerning the goodness-of-fit of a variety of disease mapping methods to simulated data for disease incidence derived from a range of models. These simulated models cover simple risk gradients to more complex true risk structures, including spatial correlation. The main general results presented here show that the gamma-Poisson exchangeable model and the Besag, York and Mollie (BYM) model are most robust across a range of diverse models. Mixture models are less robust. Non-parametric smoothing methods perform badly in general. Linear Bayes methods display behaviour similar to that of the gamma-Poisson methods.
A linear mixed effects (LME) model previously used for a spatial analysis of mortality data for a single time period is extended to include time trends and spatio-temporal interactions. This model includes functions of age and time period that can account for increasing and decreasing death rates over time and age, and a change-point of rates at a predetermined age. A geographic hierarchy is included that provides both regional and small area age-specific rate estimates, stabilizing rates based on small numbers of deaths by sharing information within a region. The proposed log-linear analysis of rates allows the use of commercially available software for parameter estimation, and provides an estimator of overdispersion directly as the residual variance. Because of concerns about the accuracy of small area rate estimates when there are many instances of no observed deaths, we consider potential sources of error, focusing particularly on the similarity of likelihood inferences using the LME model for rates as compared to an exact Poisson-normal mixed effects model for counts. The proposed LME model is applied to breast cancer deaths which occurred among white women during 1979-1996. For this example, application of diagnostics for multiparameter likelihood comparisons suggests a restriction of age to a minimum of either 25 or 35, depending on whether small area rate estimates are required. Investigation into a convergence problem led to the discovery that the changes in breast cancer geographic patterns over time are related more to urbanization than to region, as previously thought. Published in 2000 by John Wiley & Sons, Ltd.
Bayes and empirical Bayes methods have proven effective in smoothing crude maps of disease risk, eliminating the instability of estimates in low-population areas while maintaining overall geographic trends and patterns. Recent work extends these methods to the analysis of areal data which are spatially misaligned, that is, involving variables (typically counts or rates) which are aggregated over differing sets of regional boundaries. The addition of a temporal aspect complicates matters further, since now the misalignment can arise either within a given time point, or across time points (as when the regional boundaries themselves evolve over time). Hierarchical Bayesian methods (implemented via modern Markov chain Monte Carlo computing methods) enable the fitting of such models, but a formal comparison of their fit is hampered by their large size and often improper prior specifications. In this paper, we accomplish this comparison using the deviance information criterion (DIC), a recently proposed generalization of the Akaike information criterion (AIC) designed for complex hierarchical model settings like ours. We investigate the use of the delta method for obtaining an approximate variance estimate for DIC, in order to attach significance to apparent differences between models. We illustrate our approach using a spatially misaligned data set relating a measure of traffic density to paediatric asthma hospitalizations in San Diego County, California.
Maps of regional morbidity and mortality rates play an important role in assessing environmental equity. They provide effective tools for identifying areas with potentially elevated risk, determining spatial trend, and formulating and validating aetiological hypotheses about disease. Bayes and empirical Bayes methods produce stable small-area rate estimates that retain geographic and demographic resolution. The beauty of the Bayesian approach lies in its ability to structure complicated models, inferential goals and analyses. Three inferential goals are relevant to disease mapping and risk assessment: (i) computing accurate estimates of disease rates in small geographic areas; (ii) estimating the distribution of disease rates over the region; (iii) ranking the disease rates so that environmental investigation can be prioritized. No single set of estimates can simultaneously optimize these three goals, and Shen and Louis propose a set of estimates that perform well on all three goals. These are optimal for estimating the distribution of rates and for ranking, and maintain a high accuracy in estimating area-specific rates. However, the Shen/Louis method is sensitive to choice of priors. To address this issue we introduce a robustified version of the method based on a smoothed non-parametric estimate of the prior. We evaluate the performance of this method through a simulation study, and illustrate it using a data set of county-specific lung cancer rates in Ohio.
Maps of disease rates (and other quantities) often must contend with variance associated with variable population sizes and low incidence within spatial units. These characteristics can lead to substantial statistical noise that can mask underlying spatial variation. As Gelman and Price illustrated, most conventional mapping methods fail to address this problem, and in fact can introduce statistical artefacts; mapped quantities can show spatial patterns even when there are no spatial patterns in the underlying parameter of interest. Kafadar evaluated the performance of the headbanging algorithm for spatial smoothing (Tukey and Tukey, Hansen) for eliminating small scale variation and preserving edge structure. Here we perform a simulation study to investigate the artefacts of maps smoothed by unweighted and weighted headbanging. We find substantial artefacts that depend on the spatial structure of the statistical variation (for example, the spatial pattern of sample sizes) and on the details of the spatial distribution of geographic units. The methods used here could readily be adapted to study other spatial smoothers; we choose headbanging because (i) it is an important method used in practice, and (ii) its heavily computational nature is naturally studied using simulation (in contrast to the analytical methods used by Gelman and Price).
This paper aims to enlarge the usual scope of disease mapping by means of dynamic mixtures (DMDM) in case a time component is involved in the data. A special mixture model is suggested which looks for space-time components (clusters) simultaneously. The idea is illustrated using data on female lung cancer from the East German cancer registry for 1960-1989. The conventional mixed Poisson regression model is used as a third model for comparison. The models are discussed in terms of their benefits, difficulties and ease in interpretation, as well as their statistical meaning. Some ideas on evaluation of these models are also included.
A Bayesian hierarchical spatial model is constructed to describe the regional incidence of insulin dependent diabetes mellitus (IDDM) among the under 15-year-olds in Finland. The model exploits aggregated pixel-wise locations for both the cases and the population at risk. Typically such data arise from combining geographic information systems (GIS) with large databases. The dates of diagnosis and locations of the cases are observed from 1987 to 1996. The population at risk counts are available for every second year during the same period. A hierarchical model is suggested for the pixel wise case counts, including a population model to account for the uncertainty of the population at risk over the years. The model is applied in the construction of disease maps (aggregated 100 km(2) pixels), and spatial posterior predictive distributions are computed to study whether there can be found a statistically exceptional number of cases in a small area of interest.
The spatial modelling of small area health data has, for some time, included spatial autocorrelation as a random effect. This effect is non-specific and global and does not address the location of clusters of disease (a specific task). This paper addresses the need for specific and non-specific random effects within spatial epidemiology. In addition, individual frailty is also considered important and a computational algorithm based on reversible jump Markov chain Monte Carlo (RJMCMC) methods is described.
Disease incidence or disease mortality rates for small areas are often displayed on maps. Maps of raw rates, disease counts divided by the total population at risk, have been criticized as unreliable due to non-constant variance associated with heterogeneity in base population size. This has led to the use of model-based Bayes or empirical Bayes point estimates for map creation. Because the maps have important epidemiological and political consequences, for example, they are often used to identify small areas with unusually high or low unexplained risk, it is important that the assumptions of the underlying models be scrutinized. We review the use of posterior predictive model checks, which compare features of the observed data to the same features of replicate data generated under the model, for assessing model fitness. One crucial issue is whether extrema are potentially important epidemiological findings or merely evidence of poor model fit. We propose the use of the cross-validation posterior predictive distribution, obtained by reanalyzing the data without a suspect small area, as a method for assessing whether the observed count in the area is consistent with the model. Because it may not be feasible to actually reanalyze the data for each suspect small area in large data sets, two methods for approximating the cross-validation posterior predictive distribution are described.
Spatial filters have been used as an easy and intuitive way to create smoothed disease maps. Birth weight data from New York State for 1994 and 1995 are used to compare the traditional filter type of fixed geographical size with a filter size of constant or nearly constant population size. The latter are more appropriate for mapping disease in geographic areas with widely varying population density, such as New York State. Issues such as the choice of population size for the filter, the scale of smoothing, the ability to detect true spatial variation and the ability to smooth over random spatial noise are evaluated and discussed.
This paper discusses a variety of conditional autoregressive (CAR) models for mapping disease rates, beyond the usual first-order intrinsic CAR model. We illustrate the utility and scope of such models for handling different types of data structures. To encourage their routine use for map production at statistical and health agencies, a simple algorithm for fitting such models is presented. This is derived from penalized quasi-likelihood (PQL) inference which uses an analogue of best-linear unbiased estimation for the regional risk ratios and restricted maximum likelihood for the variance components. We offer the practitioner here the use of the parametric bootstrap for inference. It is more reliable than standard maximum likelihood asymptotics for inference purposes since relevant hypotheses for the mapping of rates lie on the boundary of the parameter space. We illustrate the parametric bootstrap test of the practically relevant and important simplifying hypothesis that there is no spatial autocorrelation. Although the parametric bootstrap requires computational effort, it is straightforward to implement and offers a wealth of information relating to the estimators and their properties. The proposed methodology is illustrated by analysing infant mortality in the province of British Columbia in Canada.
In this paper we discuss a number of issues that are pertinent to the analysis of disease mapping data. As an illustrative example we consider the mapping of larynx cancer across electoral wards in the North West Thames region of the U.K. Bayesian hierarchical models are now frequently employed to carry out such mapping. In a typical situation, a three-stage hierarchical model is specified in which the data are modelled as a function of area-specific relative risks at stage one; the collection of relative risks across the study region are modelled at stage two; and at stage three prior distributions are assigned to parameters of the stage two distribution. Such models allow area-specific disease relative risks to be 'smoothed' towards global and/or local mean levels across the study region. However, these models contain many structural and functional assumptions at different levels of the hierarchy; we aim to discuss some of these assumptions and illustrate their sensitivity. When relative risks are the endpoint of interest, it is common practice to assume that, for each of the age-sex strata of a particular area, there is a common multiplier (the relative risk) acting upon each of the stratum-specific risks in that area; we will examine this proportionality assumption. We also consider the choices of models and priors at stages two and three of the hierarchy, the effect of outlying areas, and an assessment of the level of smoothing that is being carried out. For inference, we concentrate on the description of the spatial variability in relative risks and on the association between the relative risks of larynx cancer and an area-level measure of socio-economic status.