Stopping rules, Bayesian reconstructions and sieves.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
We describe a probabilistic approach to simultaneous image segmentation and intensity estimation for complementary DNA microarray experiments. The approach overcomes several limitations of existing methods. In particular, it (a) uses a flexible Markov random field approach to segmentation that allows for a wider range of spot shapes than existing methods, including relatively common 'doughnut-shaped' spots; (b) models the image directly as background plus hybridization intensity, and estimates the two quantities simultaneously, avoiding the common logical error that estimates of foreground may be less than those of the corresponding background if the two are estimated separately; and (c) uses a probabilistic modeling approach to simultaneously perform segmentation and intensity estimation, and to compute spot quality measures. We describe two approaches to parameter estimation: a fast algorithm, based on the expectation-maximization and the iterated conditional modes algorithms, and a fully Bayesian framework. These approaches produce comparable results, and both appear to offer some advantages over other methods. We use an HIV experiment to compare our approach to two commercial software products: Spot and Arrayvision.
OBJECTIVE: Region-specific maps of cancer incidence, mortality, late detection rates, and screening rates can be very helpful in the planning, targeting, and coordination of cancer control activities. Unfortunately, past efforts in this area have been few, and have not used appropriate statistical models that account for the correlation of rates across both neighboring regions and different cancer types. In this article we develop such models, and apply them to the problem of cancer control in the counties of Minnesota during the period 1993-1997. METHODS: We use hierarchical Bayesian spatial statistical methods, implemented using modern Markov chain Monte Carlo computing techniques and software. RESULTS: Our approach results in spatially smoothed maps emphasizing either cancer prevention or cancer outcome for breast, colorectal, and lung cancer, as well as an overall map which combines results from these three individual cancers. CONCLUSIONS: Our methods enable us to produce a more statistically accurate picture of the geographic distribution of important cancer prevention and outcome variables in Minnesota, and appear useful for making decisions regarding targeting cancer control resources within the state.
The complexity of the global organization and internal structure of motifs in higher eukaryotic organisms raises significant challenges for motif detection techniques. To achieve successful de novo motif detection, it is necessary to model the complex dependencies within and among motifs and to incorporate biological prior knowledge. In this paper, we present LOGOS, an integrated LOcal and GlObal motif Sequence model for biopolymer sequences, which provides a principled framework for developing, modularizing, extending and computing expressive motif models for complex biopolymer sequence analysis. LOGOS consists of two interacting submodels: HMDM, a local alignment model capturing biological prior knowledge and positional dependency within the motif local structure; and HMM, a global motif distribution model modeling frequencies and dependencies of motif occurrences. Model parameters can be fit using training motifs within an empirical Bayesian framework. A variational EM algorithm is developed for de novo motif detection. LOGOS improves over existing models that ignore biological priors and dependencies in motif structures and motif occurrences, and demonstrates superior performance on both semi-realistic test data and cis-regulatory sequences from yeast and Drosophila genomes with regard to sensitivity, specificity, flexibility and extensibility.
This paper deals with the computational analysis of musical audio from recorded audio waveforms. This general problem includes, as subtasks, music transcription, extraction of musical pitch, dynamics, timbre, instrument identity, and source separation. Analysis of real musical signals is a highly ill-posed task which is made complicated by the presence of transient sounds, background interference, or the complex structure of musical pitches in the time-frequency domain. This paper focuses on models and algorithms for computer transcription of multiple musical pitches in audio, elaborated from previous work by two of the authors. The audio data are supposedly presegmented into fixed pitch regimes such as individual chords. The models presented apply to pitched (tonal) music and are formulated via a Gabor representation of nonstationary signals. A Bayesian probabilistic structure is employed for representation of prior information about the parameters of the notes. This paper introduces a numerical Bayesian inference strategy for estimation of the pitches and other parameters of the waveform. The improved algorithm is much quicker and makes the approach feasible in realistic situations. Results are presented for estimation of a known number of notes present in randomly generated note clusters from a real musical instrument database.
Standard methods for analysing survival data with covariates rely on asymptotic inferences. Bayesian methods can be performed using simple computations and are applicable for any sample size. We propose a practical method for making prior specifications and discuss a complete Bayesian analysis for parametric accelerated failure time regression models. We emphasize inferences for the survival curve rather than regression coefficients. A key feature of the Bayesian framework is that model comparisons for various choices of baseline distribution are easily handled by the calculation of Bayes factors. Such comparisons between non-nested models are difficult in the frequentist setting. We illustrate diagnostic tools and examine the sensitivity of the Bayesian methods.
The effect on intersection crashes of converting one-way street intersections in Philadelphia from signal to multiway stop sign control was estimated. Using crash and traffic volume data for a comparison group, regression models were computed to represent the normal crash experience of signal controlled intersections of one-way streets, by impact type, as a function of traffic volume. An empirical Bayesian procedure was used to estimate what would have been the expected number of crashes at the converted intersections had they not been converted. The empirical Bayesian estimates were compared with actual counts of crashes after conversion. Estimates were obtained for different classes of crashes categorized by impact type, day/night condition, and impact severity. Aggregate results indicate that replacing signals by multiway stop signs on one-way streets is associated with a reduction in crashes of approximately 24%, combining all severities, light conditions, and impact types.
Many experimental or observational studies in toxicology are best analysed in a population framework. Recent examples include investigations of the extent and origin of intra-individual variability in toxicity studies, incorporation of genotypic information to address intra-individual variability, optimal design of experiments, and extension of toxicokinetic modelling to the analysis of biomarker studies. Bayesian statistics provide powerful numerical methods for fitting population models, particularly when complex mechanistic models are involved. Challenges and limitations to the use of population models, in terms of basic structure, computational burden, ease of implementation and data accessibility, are identified and discussed.
We develop semiparametric methods for matched case-control studies using regression splines. Three methods are developed: 1) an approximate cross-validation scheme to estimate the smoothing parameter inherent in regression splines, as well as 2) Monte Carlo expectation maximization (MCEM) and 3) Bayesian methods to fit the regression spline model. We compare the approximate cross-validation approach, MCEM, and Bayesian approaches using simulation, showing that they appear approximately equally efficient; the approximate cross-validation method is computationally the most convenient. An example from equine epidemiology that motivated the work is used to demonstrate our approaches.
This article shows that the interpretation of the random-effects models used in meta-analysis to summarize heterogeneous treatment effects can have a marked effect on the results from decision models. Sources of variation in meta-analysis include the following: random variation in outcome definition (amounting to a form of measurement error), variation between the patient groups in different trials, variation between protocols, and variation in the way a given protocol is implemented. Each of these alternatives leads to a different model for how the heterogeneity in the effect sizes previously observed might relate to the effect size(s) in a future implementation. Furthermore, these alternative models require different computations and, when the net benefits are nonlinear in the efficacy parameters, result in different expected net benefits. The authors' analysis suggests that the mean treatment effect from a random-effects meta-analysis will only seldom be an appropriate representation of the efficacy expected in a future implementation. Instead, modelers should consider either the predictive distribution of a future treatment effect, or they should assume that the future implementation will result in a distribution of treatment effects. A worked example, in a probabilistic, Bayesian posterior framework, is used to illustrate the alternative computations and to show how parameter uncertainty can be combined with variation between individuals and heterogeneity in meta-analysis.
Families in which a single male is affected with a disease which might be either X linked recessive or autosomal recessive present problems in counselling. Before female relatives can be counselled, the probabilities of each mode of inheritance must be assessed, taking into account the prior probabilities, the pedigree structure, any DNA probe data, and any carrier testing data. The widely used linkage analysis package LINKAGE can be used to do the calculation, which is much simpler than the conventional Bayesian method.
A population pharmacokinetic model of intravenously and orally administered trimethoprim in patients with acquired immunodeficiency syndrome and Pneumocystis carinii pneumonia has been made using a parametric iterative two-stage Bayesian and a nonparametric expectation maximization computer program. When good information was present in the serum level data, both methods obtained similar results. With the nonparametric expectation maximization program, the median apparent rate constant for absorption (Ka) was 1.602 hr-1, median slope (Ks) of the relationship between creatinine clearance and elimination was 0.001168 hr-1, median apparent volume of distribution (Vs) was 1.058 l/kg, and median fraction of oral dose absorbed (Fa) was 0.955. These results permit dosage individualization adjusted to body weight and renal function to achieve chosen serum level peak and trough goals. Peak goals of 9 ug/ml and trough goals of 5 ug/ml appear reasonable for most patients in this population, and should permit most to complete an effective course of therapy with a reduced risk for treatment-terminating hematologic toxicity. However, therapeutic goals should always be selected based on each patient's apparent need for the drug and the risk of toxicity that is justifiably acceptable to obtain the expected benefits of the drug.
Using graph theory, we present a theoretical basis for mapping oligogenes in the joint presence of multiple phenotypic measurements of both quantitative and qualitative types. Various statistical models proposed earlier for several traits of solely single type are special cases of the unified approach given here. Our emphasis is on the generality of the framework, without specifying explicit assumptions about a sampling design. When information about environmental factors potentially affecting the traits is available, it can be incorporated into the genetic model. We adopt the Bayesian inferential machinery due to its firm theoretical basis and its capability of handling uncertain quantities; such as unobserved model parameters, missing marker data, and even different putative genetic models, probabilistically within a single framework. It is shown here that biological hypotheses about single gene affecting simultaneously multiple traits (pleiotropy) can be intuitively imposed as parameter constraints, leading to pleiotropic models for which posterior probabilities can be calculated. Outline of the possible implementation of the Bayesian method is described using the general reversible-jump Markov chain Monte Carlo algorithm. Some future challenges and extensions are also discussed.
This article outlines the statistical developments that have taken place in emission tomography during the past decade or so. We discuss the statistical aspects of the modelling of the projection data and define the additive Poisson regression model. This leads to the use of the method of maximum likelihood as a means of estimating the underlying isotope concentration within a given region of a patient's body, and to the use of the EM algorithm to compute the reconstruction. The need for the regulation of the maximum likelihood solution is tackled using Bayesian techniques. A number of algorithms for the computation of regularized solutions are outlined. The issue of parameter estimation is discussed and some open issues are mentioned.
This technical note describes the construction of posterior probability maps that enable conditional or Bayesian inferences about regionally specific effects in neuroimaging. Posterior probability maps are images of the probability or confidence that an activation exceeds some specified threshold, given the data. Posterior probability maps (PPMs) represent a complementary alternative to statistical parametric maps (SPMs) that are used to make classical inferences. However, a key problem in Bayesian inference is the specification of appropriate priors. This problem can be finessed using empirical Bayes in which prior variances are estimated from the data, under some simple assumptions about their form. Empirical Bayes requires a hierarchical observation model, in which higher levels can be regarded as providing prior constraints on lower levels. In neuroimaging, observations of the same effect over voxels provide a natural, two-level hierarchy that enables an empirical Bayesian approach. In this note we present a brief motivation and the operational details of a simple empirical Bayesian method for computing posterior probability maps. We then compare Bayesian and classical inference through the equivalent PPMs and SPMs testing for the same effect in the same data.
A significant metric in federal mammography quality standards is the phantom image quality assessment. The present work seeks to demonstrate that automated image analyses for American College of Radiology (ACR) mammographic accreditation phantom (MAP) images may be performed by a computer with objectivity, once a human acceptance level has been established. Twelve MAP images were generated with different x-ray techniques and digitized. Nineteen medical physicists in diagnostic roles (five of which were specially trained in mammography) viewed the original film images under similar conditions and provided individual scores for each test object (fibrils, microcalcifications, and nodules). Fourier domain template matching, used for low-level processing, combined with derivative filters, for intermediate-level processing, provided translation and rotation-independent localization of the test objects in the MAP images. The visibility classification decision was modeled by a Bayesian classifer using threshold contrast. The 50% visibility contrast threshold established by the trained observers' responses were: fibrils 1.010, microcalcifications 1.156, and nodules 1.016. Using these values as an estimate of human observer performance and given the automated localization of test objects, six images were graded with the computer algorithm. In all but one instance, the algorithm scored the images the same as the diagnostic physicists. In the case where it did not, the margin of disagreement was 10% due to the fact that the human scoring did not allow for half-visible fibrils (agreement occurred for the other test objects). The implication from this is that an operator-independent, machine-based scoring of MAP images is feasible and could be used as a tool to help eliminate the effect of observer variability within the current system, given proper, consistent digitization is performed.
Modeling methods in medical diagnosis are concerned with medical information processing as it pertains to utilizing biological modeling methods to facilitate patient care. Major considerations in this particular area are (1) the classification problem related to the establishment of disease entities-the taxonomy problem, and (2) the diagnosis of diseases. Available are properties, criteria, signs, symptoms, and manifestations of diseases that have been cumulated and categorized by clinicians and researchers. The problem is to optimally utilize the information content of a sign or set of signs in the practice of patient care as pertaining to the medical diagnosis problem. Some mathematical approaches implemented to facilitate such analyses include cluster analysis, discriminant analysis, Bayesian methods, computer approaches, game theory, information theory, stochastic representations, stepwise procedures, decision analysis, and pattern recognition techniques. Each of these has been studied in depth by numerous researchers advocating computer applications in medicine. Here we discuss the scope and limitations of utilizing modeling methods as a viable approach to interpreting vast amounts of biological data collected on a single patient during an encounter. We consider the following: (1) limitations associated with modeling methodologies; (2) levels of responsibilities, ranging over logging, summarizing, reporting, monitoring, and therapy selection; and (3) operational strategies and considerations as they affect hardware logistics, the actual algorithm utilized, and implementation of these sophisticated analysis systems.
Advances in statistical learning theory have resulted in a multitude of different designs of learning machines. But which ones are implemented by brains and other biological information processors? We analyze how various abstract Bayesian learners perform on different data and argue that it is difficult to determine which learning-theoretic computation is performed by a particular organism using just its performance in learning a stationary target (learning curve). Based on the fluctuation-dissipation relation in statistical physics, we then discuss a different experimental setup that might be able to solve the problem.