Search PubMed⌕ Search

Biomedical subjects

Sudipto Banerjee

Publications and source records attributed to Sudipto Banerjee.

10 recordsLinked to original sources

Rating exposure control using Bayesian decision analysis.

A model is presented for applying Bayesian statistical techniques to the problem of determining, from the usual limited number of exposure measurements, whether the exposure profile for a similar exposure group can be considered a Category 0, 1, 2, 3, or 4 exposure. The categories were adapted from the AIHA exposure category scheme and refer to (0) negligible or trivial exposure (i.e., the true X 0.95 < or =1%OEL), (1) highly controlled (i.e., X 0.95 < or =10%OEL), (2) well controlled (i.e., X 0.95 < or =50%OEL), (3) controlled (i.e., X 0.95 < or =100%OEL), or (4) poorly controlled (i.e., X0.95 > or =1%OEL) exposures. Unlike conventional statistical methods applied to exposure data, Bayesian statistical techniques can be adapted to explicitly take into account professional judgment or other sources of information. The analysis output consists of a distribution (i.e., set) of decision probabilities: e.g., 1%, 80%, 12%, 5%, and 2% probability that the exposure profile is a Category 0, 1, 2, 3, or 4 exposure. By inspection of these decision probabilities, rather than the often difficult to interpret point estimates (e.g., the sample 95th percentile exposure) and confidence intervals, a risk manager can be better positioned to arrive at an effective (i.e., correct) and efficient decision. Bayesian decision methods are based on the concepts of prior, likelihood, and posterior distributions of decision probabilities. The prior decision distribution represents what an industrial hygienist knows about this type of operation, using professional judgment; company, industry, or trade organization experience; historical or surrogate exposure data; or exposure modeling predictions. The likelihood decision distribution represents the decision probabilities based on an analysis of only the current data. The posterior decision distribution is derived by mathematically combining the functions underlying the prior and likelihood decision distributions, and represents the final decision probabilities. Advantages of Bayesian decision analysis include: (a) decision probabilities are easier to understand by risk managers and employees; (b) prior data, professional judgment, or modeling information can be objectively incorporated into the decision-making process; (c) decisions can be made with greater certainty; (d) the decision analysis can be constrained to a more realistic "parameter space" (i.e., the range of plausible values for the true geometric mean and geometric standard deviation); and (e) fewer measurements are necessary whenever the prior distribution is well defined and the process is fairly stable. Furthermore, Bayesian decision analysis provides an obvious feedback mechanism that can be used by an industrial hygienist to improve professional judgment. For example, if the likelihood decision distribution is inconsistent with the prior decision distribution then it is likely that either a significant process change has occurred or the industrial hygienist's initial judgment was incorrect. In either case, the industrial hygienist should readjust his judgment regarding this operation.

Bayes Theorem↗

Coregionalized single- and multiresolution spatially varying growth curve modeling with application to weed growth.

Modeling of longitudinal data from agricultural experiments using growth curves helps understand conditions conducive or unconducive to crop growth. Recent advances in Geographical Information Systems (GIS) now allow geocoding of agricultural data that help understand spatial patterns. A particularly common problem is capturing spatial variation in growth patterns over the entire experimental domain. Statistical modeling in these settings can be challenging because agricultural designs are often spatially replicated, with arrays of subplots, and interest lies in capturing spatial variation at possibly different resolutions. In this article, we develop a framework for modeling spatially varying growth curves as Gaussian processes that capture associations at single and multiple resolutions. We provide Bayesian hierarchical models for this setting, where flexible parameterization enables spatial estimation and prediction of growth curves. We illustrate using data from weed growth experiments conducted in Waseca, Minnesota, that recorded growth of the weed Setaria spp. in a spatially replicated design.

Bayes Theorem↗

Modelling geographically referenced survival data with a cure fraction.

The emergence of geographical information systems and related softwares nowadays enables medical databases to incorporate the geographical information on patients, allowing studies in spatial associations. Public health administrators and researchers are often interested in detecting variation in survival patterns by region or county in order to understand the possible factors that contribute towards such spatial discrepancies. These issues have led statisticians to develop survival models that account for spatial clustering and variation. Additionally, with rapid developments in medical and health sciences, researchers increasingly encounter data sets where a substantial portion of patients are cured. Models accounting for cure in the population assist in the prognosis of potentially terminal diseases. This article proposes a Bayesian modelling framework that models spatial associations for areally referenced survival data using a general class of cure models proposed by Cooner et al. The special models we outline are alternatives to the traditional proportional hazards models and can be fitted using standard Bayesian software such as WinBUGS.

Bayes Theorem↗

Semiparametric proportional odds models for spatially correlated survival data.

The last decade has witnessed major developments in Geographical Information Systems (GIS) technology resulting in the need for statisticians to develop models that account for spatial clustering and variation. In public health settings, epidemiologists and health-care professionals are interested in discerning spatial patterns in survival data that might exist among the counties. This paper develops a Bayesian hierarchical model for capturing spatial heterogeneity within the framework of proportional odds. This is deemed more appropriate when a substantial percentage of subjects enjoy prolonged survival. We discuss the implementation issues of our models, perform comparisons among competing models and illustrate with data from the SEER (Surveillance Epidemiology and End Results) database of the National Cancer Institute, paying particular attention to the underlying spatial story.

Bayes Theorem↗

Indoor air quality in two urban elementary schools--measurements of airborne fungi, carpet allergens, CO2, temperature, and relative humidity.

This article presents measurements of biological contaminants in two elementary schools that serve inner city minority populations. One of the schools is an older building; the other is newer and was designed to minimize indoor air quality problems. Measurements were obtained for airborne fungi, carpet loadings of dust mite allergens, cockroach allergens, cat allergens, and carpet fungi. Carbon dioxide concentrations, temperature, and relative humidity were also measured. Each of these measurements was made in five classrooms in each school over three seasons--fall, winter, and spring. We compared the indoor environments at the two schools and examined the variability in measured parameters between and within schools and across seasons. A fixed-effects, nested analysis was performed to determine the effect of school, season, and room-within-school, as well as CO2, temperature and relative humidity. The levels of all measured parameters were comparable for the two schools. Carpet culturable fungal concentrations and cat allergen levels in the newer school started and remained higher than in the older school over the study period. Cockroach allergen levels in some areas were very high in the newer school and declined over the study period to levels lower than the older school. Dust mite allergen and culturable fungal concentrations in both schools were relatively low compared with benchmark values. The daily averages for temperature and relative humidity frequently did not meet ASHRAE guidelines in either school, which suggests that proper HVAC and general building operation and maintenance procedures are at least as important as proper design and construction for adequate indoor air quality. The results show that for fungi and cat allergens, the school environment can be an important exposure source for children.

Air Microbiology↗

On geodetic distance computations in spatial modeling.

Statisticians analyzing spatial data often need to detect and model associations based upon distances on the Earth's surface. Accurate computation of distances are sought for exploratory and interpretation purposes, as well as for developing numerically stable estimation algorithms. When the data come from locations on the spherical Earth, application of Euclidean or planar metrics for computing distances is not straightforward. Yet, planar metrics are desirable because of their easier interpretability, easy availability in software packages, and well-established theoretical properties. While distance computations are indispensable in spatial modeling, their importance and impact upon statistical estimation and prediction have gone largely unaddressed. This article explores the different options in using planar metrics and investigates their impact upon spatial modeling.

Algorithms↗

Generalized hierarchical multivariate CAR models for areal data.

In the fields of medicine and public health, a common application of areal data models is the study of geographical patterns of disease. When we have several measurements recorded at each spatial location (for example, information on p>/= 2 diseases from the same population groups or regions), we need to consider multivariate areal data models in order to handle the dependence among the multivariate components as well as the spatial dependence between sites. In this article, we propose a flexible new class of generalized multivariate conditionally autoregressive (GMCAR) models for areal data, and show how it enriches the MCAR class. Our approach differs from earlier ones in that it directly specifies the joint distribution for a multivariate Markov random field (MRF) through the specification of simpler conditional and marginal models. This in turn leads to a significant reduction in the computational burden in hierarchical spatial random effect modeling, where posterior summaries are computed using Markov chain Monte Carlo (MCMC). We compare our approach with existing MCAR models in the literature via simulation, using average mean square error (AMSE) and a convenient hierarchical model selection criterion, the deviance information criterion (DIC; Spiegelhalter et al., 2002, Journal of the Royal Statistical Society, Series B64, 583-639). Finally, we offer a real-data application of our proposed GMCAR approach that models lung and esophagus cancer death rates during 1991-1998 in Minnesota counties.

Bayes Theorem↗

Parametric spatial cure rate models for interval-censored time-to-relapse data.

Several recent papers (e.g., Chen, Ibrahim, and Sinha, 1999, Journal of the American Statistical Association 94, 909-919; Ibrahim, Chen, and Sinha, 2001a, Biometrics 57, 383-388) have described statistical methods for use with time-to-event data featuring a surviving fraction (i.e., a proportion of the population that never experiences the event). Such cure rate models and their multivariate generalizations are quite useful in studies of multiple diseases to which an individual may never succumb, or from which an individual may reasonably be expected to recover following treatment (e.g., various types of cancer). In this article we extend these models to allow for spatial correlation (estimable via zip code identifiers for the subjects) as well as interval censoring. Our approach is Bayesian, where posterior summaries are obtained via a hybrid Markov chain Monte Carlo algorithm. We compare across a broad collection of rather high-dimensional hierarchical models using the deviance information criterion, a tool recently developed for just this purpose. We apply our approach to the analysis of a smoking cessation study where the subjects reside in 53 southeastern Minnesota zip codes. In addition to the usual posterior estimates, our approach yields smoothed zip code level maps of model parameters related to the relapse rates over time and the ultimate proportion of quitters (the cure rates).

Algorithms↗

Expert judgment and occupational hygiene: application to aerosol speciation in the nickel primary production industry.

In many situations characterized by sparse data, occupational hygienists have used subjective judgments that are claimed to be derived from their experience and knowledge. While this practice is widespread, there has been no systematic study of 'expert judgment' or the 'art' of occupational hygiene. Indeed, there is a need to address the question of whether there is such a thing as 'expert opinion' in occupational hygiene that is broadly shared by practicing professionals. This research, employing 11 experts who estimate an exposure parameter (the percentages of four nickel species) in 12 workplaces in a nickel primary production industry, provides a large dataset from which useful inferences can be drawn about the quality of expert judgments and the variability among the experts. A well-designed questionnaire that provided succinct information about the processes and baseline data served to calibrate the experts. The Bayesian framework has been used in this work to develop posterior means and standard deviations of the percentages of the four nickel species in the 12 workplaces of interest in the company. These estimates of the nickel speciation are at least as precise as--and most of the time more precise than--those provided by the sparse measurement data. There was a very high degree of agreement among the experts. A majority of the experts agreed among themselves 92% of the time, while almost two-thirds agreed 73% of the time. This, coupled with the fact that the experts came from varied backgrounds, seems to suggest that there is indeed some broad body of specialized knowledge that the experts are drawing on to reach similar judgments. It also seems that one type of expert is not necessarily any better than any other kind, and expertise does not necessarily require intimate familiarity with the workplace. In this example, the expert judgment exercise has indeed enhanced the quality of our knowledge of the exposure 'fingerprints' for the nickel industry workplaces studied and the combination of expert judgment and sparse data is better than the sparse data alone. For occupational hygiene exposure assessment, our experience suggests that such expert judgment methods can provide a cost-effective means to improve and refine information about workplace hazards. However, more study is warranted for situations where the domain of the quantity of interest has a much wider range of values, e.g. actual exposure values.

Air Pollutants, Occupational↗

Frailty modeling for spatially correlated survival data, with application to infant mortality in Minnesota.

The use of survival models involving a random effect or 'frailty' term is becoming more common. Usually the random effects are assumed to represent different clusters, and clusters are assumed to be independent. In this paper, we consider random effects corresponding to clusters that are spatially arranged, such as clinical sites or geographical regions. That is, we might suspect that random effects corresponding to strata in closer proximity to each other might also be similar in magnitude. Such spatial arrangement of the strata can be modeled in several ways, but we group these ways into two general settings: geostatistical approaches, where we use the exact geographic locations (e.g. latitude and longitude) of the strata, and lattice approaches, where we use only the positions of the strata relative to each other (e.g. which counties neighbor which others). We compare our approaches in the context of a dataset on infant mortality in Minnesota counties between 1992 and 1996. Our main substantive goal here is to explain the pattern of infant mortality using important covariates (sex, race, birth weight, age of mother, etc.) while accounting for possible (spatially correlated) differences in hazard among the counties. We use the GIS ArcView to map resulting fitted hazard rates, to help search for possible lingering spatial correlation. The DIC criterion (Spiegelhalter et al., Journal of the Royal Statistical Society, Series B 2002, to appear) is used to choose among various competing models. We investigate the quality of fit of our chosen model, and compare its results when used to investigate neonatal versus post-neonatal mortality. We also compare use of our time-to-event outcome survival model with the simpler dichotomous outcome logistic model. Finally, we summarize our findings and suggest directions for future research.

Adult↗