Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Bayesian inference”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Multiple outputation: inference for complex clustered data by averaging analyses from independent data.

This article applies a simple method for settings where one has clustered data, but statistical methods are only available for independent data. We assume the statistical method provides us with a normally distributed estimate, theta, and an estimate of its variance sigma. We randomly select a data point from each cluster and apply our statistical method to this independent data. We repeat this multiple times, and use the average of the associated theta's as our estimate. An estimate of the variance is given by the average of the sigma2's minus the sample variance of the theta's. We call this procedure multiple outputation, as all "excess" data within each cluster is thrown out multiple times. Hoffman, Sen, and Weinberg (2001, Biometrika 88, 1121-1134) introduced this approach for generalized linear models when the cluster size is related to outcome. In this article, we demonstrate the broad applicability of the approach. Applications to angular data, p-values, vector parameters, Bayesian inference, genetics data, and random cluster sizes are discussed. In addition, asymptotic normality of estimates based on all possible outputations, as well as a finite number of outputations, is proven given weak conditions. Multiple outputation provides a simple and broadly applicable method for analyzing clustered data. It is especially suited to settings where methods for clustered data are impractical, but can also be applied generally as a quick and simple tool.

Acute Disease↗

The contributions of Jerome Cornfield to the theory of statistics.

This paper is a review of the contributions of Jerome Cornfield to the theory of statistics. It discusses several highlights of his theoretical work as well as describing his philosophy relating theory to application. The three areas discussed are: linear programming, urn sampling and its generalizations to the analysis of variance, and Bayesian inference. It is not widely known that Jerome Cornfield was perhaps the first to formulate and approximately solve the linear programming problem in 1941. His formulation was made for the famous "Diet Problem". An early publication introduced the method of indicator random variables in the context of urn sampling. This simple method allowed straightforward calculations of the low order moments for estimates arising from sampling finite populations and was later generalized to the two-way analysis of variance. The application of the urn sampling model to the analysis of variance served to illuminate how one chooses proper error terms for making tests in the analysis of variance table. Jerome Cornfield's philosophy on applications of statistics was dominated by a Bayesian outlook. His theoretical contributions in the past two decades were mainly concerned with the development of Bayesian ideas and methods. A brief survey is made of his main contributions to this area. A particularly noteworthy result was his demonstration that for the two-sample slippage problem of location, the likelihood function under a permutation setting is uninformative for the slippage parameter. However, the posterior distribution differs from the prior distribution despite the fact that the likelihood is uninformative.

Bayes Theorem↗

Bayesian analysis of prevalence with covariates using simulation-based techniques: applications to HIV screening.

Ignoring the limited precision of medical diagnostic tests can incur serious bias in prevalence estimation. Conversely, treating the values of sensitivity and specificity as constants, as in most studies, inevitably underestimates the variability of prevalence estimates. Bayesian inference provides a natural framework with which to integrate the variability in the estimates of sensitivity and specificity with estimation of prevalence. However, the resulting model becomes quite complicated and presents a computational challenge. Recently, Mendoza-Blanco et al. proposed a missing-data approach with simulation-based techniques to deal with the computational difficulties. Although their approach is quite effective in reducing the computational complexity into manageable tasks, their developed methodology is not general enough for modelling the effects of covariates in prevalence estimation. In this paper, we extend their work in this direction by combining their missing-data approach with a latent variable technique for modelling discrete data. The present work also generalizes the methods of Albert and Chib for Bayesian analysis of binary response data with errors in the response. We illustrate the methodology with several real data examples extracted from the literature.

AIDS Serodiagnosis↗

Unraveling the evolutionary radiation of the thoracican barnacles using molecular and morphological evidence: a comparison of several divergence time estimation approaches.

The Thoracica includes the ordinary barnacles found along the sea shore and is the most diverse and well-studied superorder of Cirripedia. However, although the literature abounds with scenarios explaining the evolution of these barnacles, very few studies have attempted to test these hypotheses in a phylogenetic context. The few attempts at phylogenetic analyses have suffered from a lack of phylogenetic signal and small numbers of taxa. We collected DNA sequences from the nuclear 18S, 28S, and histone H3 genes and the mitochondrial 12S and 16S genes (4,871 bp total) and data for 37 adult and 53 larval morphological characters from 43 taxa representing all the extant thoracican suborders (except the monospecific Brachylepadomorpha). Four Rhizocephala (highly modified parasitic barnacles) taxa and a Rhizocephala + Acrothoracica (burrowing barnacles) hypothetical ancestor were used as the outgroup for the molecular and morphological analyses, respectively. We analyzed these data separately and combined using maximum likelihood (ML) under "hill-climbing" and genetic algorithm heuristic searches, maximum parsimony procedures, and Bayesian inference coupled with Markov chain Monte Carlo techniques under mixed and homogeneous models of nucleotide substitution. The resulting phylogenetic trees answered key questions in barnacle evolution. The four-plated Iblomorpha were shown as the most primitive thoracican, and the plateless Heteralepadomorpha were placed as the sister group of the Lepadomorpha. These relationships suggest for the first time in an invertebrate that exoskeleton biomineralization may have evolved from phosphatic to calcitic. Sessilia (nonpedunculate) barnacles were depicted as monophyletic and appear to have evolved from a stalked (pedunculate) multiplated (5+) scalpelloidlike ancestor rather than a five-plated lepadomorphan ancestor. The Balanomorpha (symmetric sessile barnacles) appear to have the following relationship: (Chthamaloidea(Coronuloidea(Tetraclitoidea, Balanoidea))). Thoracican divergence times were estimated under ML-based local clock, Bayesian, and penalized likelihood approaches using an 18S data set and three calibration points: Heteralepadomorpha = 530 million years ago (MYA), Scalpellomorpha = 340 MYA, and Verrucomorpha = 120 MYA. Estimated dates varied considerably within and between approaches depending on the calibration point. Highly parameterized local clock models that assume independent rates (r > or = 15) for confamilial or congeneric species generated the most congruent estimates among calibrations and agreed more closely with the barnacle fossil record. Reasonable estimates were also obtained under the Bayesian procedure of Kishino et al. (2001, Mol. Biol. Evol. 18:352-361) but using multiple calibrations. Most of the dates estimated under the Bayesian procedure of Aris-Brosou and Yang (2002, Syst. Biol. 51:703-714) and the penalized likelihood method using single and/or multiple calibrations were inconsistent among calibrations and did not fit the fossil record.

Animals↗

Models of brain function in neuroimaging.

Inferences about brain function, using neuroimaging data, rest on models of how the data were caused. These models can be quite diverse, ranging from conceptual models of functional anatomy to nonlinear mathematical models of hemodynamics. However, they all have to be internally consistent because they model the same thing. This consistency encompasses many levels of description and places constraints on the statistical models, adopted for data analysis, and the experimental designs they embody. The aim of this review is to introduce the key models used in imaging neuroscience and how they relate to each other. We start with anatomical models of functional brain architectures, which motivate some of the fundaments of neuroimaging. We then turn to basic statistical models (e.g., the general linear model) used for making classical and Bayesian inferences about where neuronal responses are expressed. By incorporating biophysical constraints, these basic models can be finessed and, in a dynamic setting, rendered causal. This allows us to infer how interactions among brain regions are mediated.

Biophysical Phenomena↗

Molecular systematics and biogeography of the southern South american freshwater "crabs" Aegla (decapoda: Anomura: Aeglidae) using multiple heuristic tree search approaches.

Recently new heuristic genetic algorithms such as Treefinder and MetaGA have been developed to search for optimal trees in a maximum likelihood (ML) framework. In this study we combined these methods with other standard heuristic approaches such as ML and maximum parsimony hill-climbing searches and Bayesian inference coupled with Markov chain Monte Carlo techniques under homogeneous and mixed models of evolution to conduct an extensive phylogenetic analysis of the most abundant and widely distributed southern South American freshwater"crab,"the Aegla(Anomura: Aeglidae). A total of 167 samples representing 64 Aegla species and subspecies were sequenced for one nuclear (28S rDNA) and four mitochondrial (12S and 16S rDNA, COI, and COII) genes (5352 bp total). Additionally, six other anomuran species from the genera Munida,Pachycheles, and Uroptychus(Galatheoidea), Lithodes(Paguroidea), and Lomis(Lomisoidea) and the nuclear 18S rDNA gene (1964 bp) were included in preliminary analyses for rooting the Aegla tree. Nonsignificantly different phylogenetic hypotheses resulted from all the different heuristic methods used here, although the best scored topologies found under the ML hill-climbing, Bayesian, and MetaGA approaches showed considerably better likelihood scores (Delta> 54) than those found under the MP and Treefinder approaches. Our trees provided strong support for most of the recognized Aegla species except for A. cholchol,A. jarai,A. parana,A. marginata, A. platensis, and A. franciscana, which may actually represent multiple species. Geographically, the Aegla group was divided into a basal western clade (21 species and subspecies) composed of two subclades with overlapping distributions, and a more recent central-eastern clade (43 species) composed of three subclades with fairly well-recognized distributions. This result supports the Pacific-Origin Hypothesis postulated for the group; alternative hypotheses of Atlantic or multiple origins were significantly rejected by our analyses. Finally, we combined our phylogenetic results with previous hypotheses of South American paleodrainages since the Jurassic to propose a biogeographical framework of the Aegla radiation.

Animals↗

Singularities affect dynamics of learning in neuromanifolds.

The parameter spaces of hierarchical systems such as multilayer perceptrons include singularities due to the symmetry and degeneration of hidden units. A parameter space forms a geometrical manifold, called the neuromanifold in the case of neural networks. Such a model is identified with a statistical model, and a Riemannian metric is given by the Fisher information matrix. However, the matrix degenerates at singularities. Such a singular structure is ubiquitous not only in multilayer perceptrons but also in the gaussian mixture probability densities, ARMA time-series model, and many other cases. The standard statistical paradigm of the Cramér-Rao theorem does not hold, and the singularity gives rise to strange behaviors in parameter estimation, hypothesis testing, Bayesian inference, model selection, and in particular, the dynamics of learning from examples. Prevailing theories so far have not paid much attention to the problem caused by singularity, relying only on ordinary statistical theories developed for regular (nonsingular) models. Only recently have researchers remarked on the effects of singularity, and theories are now being developed. This article gives an overview of the phenomena caused by the singularities of statistical manifolds related to multilayer perceptrons and gaussian mixtures. We demonstrate our recent results on these problems. Simple toy models are also used to show explicit solutions. We explain that the maximum likelihood estimator is no longer subject to the gaussian distribution even asymptotically, because the Fisher information matrix degenerates, that the model selection criteria such as AIC, BIC, and MDL fail to hold in these models, that a smooth Bayesian prior becomes singular in such models, and that the trajectories of dynamics of learning are strongly affected by the singularity, causing plateaus or slow manifolds in the parameter space. The natural gradient method is shown to perform well because it takes the singular geometrical structure into account. The generalization error and the training error are studied in some examples.

Journal Article↗

Trial-to-trial variability of cortical evoked responses: implications for the analysis of functional connectivity.

OBJECTIVES: The time series of single trial cortical evoked potentials typically have a random appearance, and their trial-to-trial variability is commonly explained by a model in which random ongoing background noise activity is linearly combined with a stereotyped evoked response. In this paper, we demonstrate that more realistic models, incorporating amplitude and latency variability of the evoked response itself, can explain statistical properties of cortical potentials that have often been attributed to stimulus-related changes in functional connectivity or other intrinsic neural parameters. METHODS: Implications of trial-to-trial evoked potential variability for variance, power spectrum, and interdependence measures like cross-correlation and spectral coherence, are first derived analytically. These implications are then illustrated using model simulations and verified experimentally by the analysis of intracortical local field potentials recorded from monkeys performing a visual pattern discrimination task. To further investigate the effects of trial-to-trial variability on the aforementioned statistical measures, a Bayesian inference technique is used to separate single-trial evoked responses from the ongoing background activity. RESULTS: We show that, when the average event-related potential (AERP) is subtracted from single-trial local field potential time series, a stimulus phase-locked component remains in the residual time series, in stark contrast to the assumption of the common model that no such phase-locked component should exist. Two main consequences of this observation are demonstrated for statistical measures that are computed on the residual time series. First, even though the AERP has been subtracted, the power spectral density, computed as a function of time with a short sliding window, can nonetheless show signs of modulation by the AERP waveform. Second, if the residual time series of two channels co-vary, then their cross-correlation and spectral coherence time functions can also be modulated according to the shape of the AERP waveform. Bayesian estimation of single-trial evoked responses provides further proof that these time-dependent statistical changes are due to remnants of the evoked phase-locked component in the residual time series. CONCLUSIONS: Because trial-to-trial variability of the evoked response is commonly ignored as a contributing factor in evoked potential studies, stimulus-related modulations of power spectral density, cross-correlation, and spectral coherence measures is often attributed to dynamic changes of the connectivity within and among neural populations. This work demonstrates that trial-to-trial variability of the evoked response must be considered as a possible explanation of such modulation.

Animals↗

Plastome evolution and phylogenomic relationships in Ajuga (Lamiaceae, Ajugoideae).

BACKGROUND: Ajuga is currently known to include approximately 69 species, with a combined distribution extending throughout Eurasia, Africa, and Australia. Its popularity and significance are largely based on an extensive history of medicinal and horticultural use. It is divided into two sections based on morphological characters, and this sectional classification is also reflected in pronounced geographic patterns. Although previous studies have largely focused on Ajuga sect. Ajuga in East Asia, A. sect. Chamaepithys, which ranges from the Mediterranean to Central Asia, remains insufficiently sampled, thereby limiting a comprehensive understanding of infrageneric sectional relationships within the genus. Here, we generated complete plastid genomes for 12 species representing both sections of the genus and used these data to characterize plastome structure and infer evolutionary relationships. RESULTS: In this study, 21 Ajuga plastomes were analyzed, including 12 newly sequenced plastomes and 9 previously published plastomes representing 19 species. Comparative analyses showed that all plastomes exhibited a highly conserved quadripartite structure, with genome sizes ranging from 149,963 to 150,740 bp and GC contents varying from 38.2% to 38.3%. Each plastome contained 133 genes, including 88 protein-coding genes, 37 transfer RNA genes, and 8 ribosomal RNA genes. The boundaries between the inverted repeat (IR) and single-copy (SC) regions were also highly conserved across species. In addition, 796 simple sequence repeats (SSRs), 874 long repeat sequences (LRSs), and 12 highly variable regions (ccsA-ndhD, ndhF-rpl32, petA-psbJ, rpl32-trnL-UAG, rps2-rpoC2, trnH-GUG-psbA, trnK-UUU-rps16, trnP-UGG-psaJ, trnT-UGU-trnL-UAA, ycf15-trnL-CAA, ndhF, and ycf1) were identified among the 21 plastomes. Phylogenetic analyses based on four datasets and conducted using Maximum Likelihood and Bayesian Inference recovered two major clades corresponding to the traditionally recognized sectional classification, with one distributed from the Mediterranean to Central Asia and the other in East Asia. CONCLUSION: This study represents the most comprehensive plastome-based sampling of Ajuga to date, including representative species from the Mediterranean, Central Asia, and East Asia. Our results have significantly enhanced our understanding of its infrageneric relationships. The plastome resources generated in this study provide a valuable foundation for future research on species delimitation, phylogeny, and the evolutionary history of Ajuga.

Phylogeny↗

Evolutionary HMMs: a Bayesian approach to multiple alignment.

MOTIVATION: We review proposed syntheses of probabilistic sequence alignment, profiling and phylogeny. We develop a multiple alignment algorithm for Bayesian inference in the links model proposed by Thorne et al. (1991, J. Mol. Evol., 33, 114-124). The algorithm, described in detail in Section 3, samples from and/or maximizes the posterior distribution over multiple alignments for any number of DNA or protein sequences, conditioned on a phylogenetic tree. The individual sampling and maximization steps of the algorithm require no more computational resources than pairwise alignment. METHODS: We present a software implementation (Handel) of our algorithm and report test results on (i) simulated data sets and (ii) the structurally informed protein alignments of BAliBASE (Thompson et al., 1999, Nucleic Acids Res., 27, 2682-2690). RESULTS: We find that the mean sum-of-pairs score (a measure of residue-pair correspondence) for the BAliBASE alignments is only 13% lower for Handelthan for CLUSTALW(Thompson et al., 1994, Nucleic Acids Res., 22, 4673-4680), despite the relative simplicity of the links model (CLUSTALW uses affine gap scores and increased penalties for indels in hydrophobic regions). With reference to these benchmarks, we discuss potential improvements to the links model and implications for Bayesian multiple alignment and phylogenetic profiling. AVAILABILITY: The source code to Handelis freely distributed on the Internet at http://www.biowiki.org/Handel under the terms of the GNU Public License (GPL, 2000, http://www.fsf.org./copyleft/gpl.html).

Algorithms↗

A molecular supermatrix of the rabbits and hares (Leporidae) allows for the identification of five intercontinental exchanges during the Miocene.

The hares and rabbits belonging to the family Leporidae have a nearly worldwide distribution and approximately 72% of the genera have geographically restricted distributions. Despite several attempts using morphological, cytogenetic, and mitochondrial DNA evidence, a robust phylogeny for the Leporidae remains elusive. To provide phylogenetic resolution within this group, a molecular supermatrix was constructed for 27 taxa representing all 11 leporid genera. Five nuclear (SPTBN1, PRKCI, THY, TG, and MGF) and two mitochondrial (cytochrome b and 12S rRNA) gene fragments were analyzed singly and in combination using parsimony, maximum likelihood, and Bayesian inference. The analysis of each gene fragment separately as well as the combined mtDNA data almost invariably failed to provide strong statistical support for intergeneric relationships. In contrast, the combined nuclear DNA topology based on 3601 characters greatly increased phylogenetic resolution among leporid genera, as was evidenced by the number of topologies in the 95% confidence interval and the number of significantly supported nodes. The final molecular supermatrix contained 5483 genetic characters and analysis thereof consistently recovered the same topology across a range of six arbitrarily chosen model specifications. Twelve unique insertion-deletions were scored and all could be mapped to the tree to provide additional support without introducing any homoplasy. Dispersal-vicariance analyses suggest that the most parsimonious solution explaining the current geographic distribution of the group involves an Asian or North American origin for the Leporids followed by at least nine dispersals and five vicariance events. Of these dispersals, at least three intercontinental exchanges occurred between North America and Asia via the Bering Strait and an additional three independent dispersals into Africa could be identified. A relaxed Bayesian molecular clock applied to the seven loci used in this study indicated that most of the intercontinental exchanges occurred between 14 and 9 million years ago and this period is broadly coincidental with the onset of major Antarctic expansions causing land bridges to be exposed.

Animals↗

The complete sequences and gene organisation of the mitochondrial genomes of the heterodont bivalves Acanthocardia tuberculata and Hiatella arctica--and the first record for a putative Atpase subunit 8 gene in marine bivalves.

BACKGROUND: Mitochondrial (mt) gene arrangement is highly variable among molluscs and especially among bivalves. Of the 30 complete molluscan mt-genomes published to date, only one is of a heterodont bivalve, although this is the most diverse taxon in terms of species numbers. We determined the complete sequence of the mitochondrial genomes of Acanthocardia tuberculata and Hiatella arctica, (Mollusca, Bivalvia, Heterodonta) and describe their gene contents and genome organisations to assess the variability of these features among the Bivalvia and their value for phylogenetic inference. RESULTS: The size of the mt-genome in Acanthocardia tuberculata is 16.104 basepairs (bp), and in Hiatella arctica 18.244 bp. The Acanthocardia mt-genome contains 12 of the typical protein coding genes, lacking the Atpase subunit 8 (atp8) gene, as all published marine bivalves. In contrast, a complete atp8 gene is present in Hiatella arctica. In addition, we found a putative truncated atp8 gene when re-annotating the mt-genome of Venerupis philippinarum. Both mt-genomes reported here encode all genes on the same strand and have an additional trnM. In Acanthocardia several large non-coding regions are present. One of these contains 3.5 nearly identical copies of a 167 bp motive. In Hiatella, the 3' end of the NADH dehydrogenase subunit (nad)6 gene is duplicated together with the adjacent non-coding region. The gene arrangement of Hiatella is markedly different from all other known molluscan mt-genomes, that of Acanthocardia shows few identities with the Venerupis philippinarum. Phylogenetic analyses on amino acid and nucleotide levels robustly support the Heterodonta and the sister group relationship of Acanthocardia and Venerupis. Monophyletic Bivalvia are resolved only by a Bayesian inference of the nucleotide data set. In all other analyses the two unionid species, being to only ones with genes located on both strands, do not group with the remaining bivalves. CONCLUSION: The two mt-genomes reported here add to and underline the high variability of gene order and presence of duplications in bivalve and molluscan taxa. Some genomic traits like the loss of the atp8 gene or the encoding of all genes on the same strand are homoplastic among the Bivalvia. These characters, gene order, and the nucleotide sequence data show considerable potential of resolving phylogenetic patterns at lower taxonomic levels.

Journal Article↗

Estimation of population pharmacokinetic parameters in the presence of non-compliance.

In population pharmacokinetic (PK) studies, patients' drug plasma profiles are routinely analyzed assuming that all patients took their drug at the times and in the amounts specified. However, patient non-compliance with the prescribed drug regimen is a leading source of failure to drug therapy. It has been reported that over 30% of patients routinely skip doses regardless of their disease, prognosis, or symptoms. This brings into question the assumption regarding full compliance for population PK analyses. This paper describes the estimation of population PK parameters in the presence and absence of non-compliance while either assuming full compliance or estimating compliance using a hierarchical Bayesian approach. Assessment of compliance for a given dose was limited to one of three possibilities: no dose was taken at the prescribed time, the prescribed dose was taken at the prescribed time, or twice the prescribed dose was taken at the prescribed time. Simulated data sets based on a one-compartment pharmacokinetic model with first order elimination were analyzed using WinBUGS (Bayesian inference Using Gibbs Sampling) software. An initial feasibility simulation experiment, using a simple, but informative PK sampling design with bolus input of drug, was performed. A second simulation study was then carried out using a more realistic sampling design and first-order input of drug. The simulated sampling design included observations after known doses as well as after uncertain doses. Results from the feasibility study revealed that when compliance was estimated instead of being assumed to be 100%, the relative prediction error for clearance (CL) decreased from 0.25 to 0.10 for 60% compliance and from 0.6 to 0.2 for 35% compliance. Estimates of the interoccasion variability of clearance were improved by compliance estimation but still had substantial positive bias. Estimated of interindividual variability were relatively insensitive to compliance estimation. Estimates for volume of distribution (V) and its associated variances were not affected by incorporation of compliance estimates, perhaps due to the specific sampling design that was used. The design was relatively uninformative regarding V. In the more realistic study, estimates for CL, V and the difference between the absorption rate constant and the elimination rate constant (KA-K) were improved by the incorporation of compliance estimation. The median relative errors were reduced from 0.51 to -0.01 for CL, from 0.49 to 0.04 for V, and from 0.49 to -0.02 for Ka-K. The bias in interoccasion variances for V and CL appeared to be reduced by compliance estimation while estimates of interindividual variability were not affected in a systematic fashion. The bias in the residual error variance was decreased from a relative error of about 2 to close to 0. The use of hierarchical Bayesian modeling with the incorporation of compliance estimation decreased the bias in the typical value parameter but the effects on variance parameters were less consistent. The encouraging results of these simulation experiments will hopefully stimulate further evaluation of this methodology for the estimation of population pharmacokinetic parameters in the presence of potential patient noncompliance.

Bayes Theorem↗

Bayesian analysis for a single 2 x 2 table.

The simple comparison of two binomial populations is frequently of interest in epidemiology when the domains are large. For small domains, however, there are no exact methods except Fisher's exact test. A basic problem, therefore, is to compare two populations by assessing the difference between the proportions of individuals who possess a characteristic in the first and second populations. When there is prior information, we take the proportions to have independent conjugate beta distributions with known parameters, thereby facilitating a Bayesian analysis. We consider Bayesian inference on functions of the proportions, and the three most common scalar measures used in epidemiology and health services research, namely relative risk, odds ratio and attributable risk. We develop the highest density regions (both exact and approximate) for relative risk, odds ratio and attributable risk. In addition, we consider the Bayes factor for testing whether the model with a common proportion holds rather than one with distinct proportions. Using data from the population-based Worcester Heart Attack Study, we apply our methodology to study gender differences in the therapeutic management of patients with acute myocardial infarction (AMI) by selected demographic and clinical characteristics. The Bayes factor, the approximate and exact intervals generally suggest that there are no substantial differences in the pharmacologic management of males and females hospitalized with AMI.

Adult↗

Identifying the types of missingness in quality of life data from clinical trials.

This paper discusses methods of identifying the types of missingness in quality of life (QOL) data in cancer clinical trials. The first approach involves collecting information on why the QOL questionnaires were not completed. Based on the reasons provided one may be able to distinguish the mechanisms causing missing data. The second approach is to model the missing data mechanism and perform hypothesis testing to determine the missing data processes. Two methods of testing if missing data are missing completely at random (MCAR) are presented and applied to incomplete longitudinal QOL data obtained from international multi-centre cancer clinical trials. The first method (Ridout, 1991) is based on a logistic regression and the second method (Park and Davis, 1993) is based on an adaptation of weighted least squares. In one application (advanced breast cancer) missing data was not likely to be MCAR. In the second application (adjuvant breast cancer) the missing mechanism was dependent on the QOL scale under study. MCAR and missing at random (MAR) have distinct consequences for data analysis. Therefore it is relevant to distinguish between them. However, if either MCAR or MAR hold, likelihood or Bayesian inferences can be based solely on the observed data, although for MAR, depending on the research question, modelling the dropout mechanism may still be necessary. Distinguishing between MAR and missing not at random (MNAR) is not trivial and relies on fundamentally untestable assumptions.

Clinical Trials as Topic↗

Assessing heterogeneity and correlation of paired failure times with the bivariate frailty model.

We consider bivariate survival times for heterogeneous populations, where heterogeneity induces deviations in an individual's risk of an event as well as associations between survival times. The heterogeneity is characterized by a bivariate frailty model. We measure the heterogeneity effects through deviations associated with hazard functions and an association function defined through the conditional hazard functions: the cross-ratio function proposed by Oakes. We show how the deviation and association measures are determined by the frailty distribution. A Gibbs sampling method is developed for Bayesian inferences on regression coefficients, frailty parameters and the heterogeneity measures. The method is applied to a mental health care data set.

Algorithms↗

A bayesian analysis for spatial processes with application to disease mapping.

In epidemiology, maps of disease rates and disease risk provide a spatial perspective for researching disease aetiology. For rare diseases or when the population base is small, the rate and risk estimates may be unstable. We propose using a Bayesian analysis based on the conditional autoregressive (CAR) process that will spatially smooth disease rates or risk estimates by allowing each site to 'borrow strength' from its neighbours. Covariates may be included in the model in such a way as to establish a possible association between risk factors and disease incidence. Bayesian inferences are implemented from a direct resampling scheme where large samples are generated from the various posterior distributions. The methodology is demonstrated with a simulation that assesses the effect of sample size and the model parameters on inferences for the parameters. Our approach is also used to spatially smooth district lip cancer rates in Scotland using the CAR model with a covariate that allows for exposure to sunlight.

Bayes Theorem↗

Genetic variance components analysis for binary phenotypes using generalized linear mixed models (GLMMs) and Gibbs sampling.

The common complex diseases such as asthma are an important focus of genetic research, and studies based on large numbers of simple pedigrees ascertained from population-based sampling frames are becoming commonplace. Many of the genetic and environmental factors causing these diseases are unknown and there is often a strong residual covariance between relatives even after all known determinants are taken into account. This must be modelled correctly whether scientific interest is focused on fixed effects, as in an association analysis, or on the covariances themselves. Analysis is straightforward for multivariate Normal phenotypes, but difficulties arise with other types of trait. Generalized linear mixed models (GLMMs) offer a potentially unifying approach to analysis for many classes of phenotype including multivariate Normal traits, binary traits, and censored survival times. Markov Chain Monte Carlo methods, including Gibbs sampling, provide a convenient framework within which such models may be fitted. In this paper, Bayesian inference Using Gibbs Sampling (a generic Gibbs sampler; BUGS) is used to fit GLMMs for multivariate Normal and binary phenotypes in nuclear families. BUGS is easy to use and readily available. We motivate a suitable model structure for Normal phenotypes and show how the model extends to binary traits. We discuss parameter interpretation and statistical inference and show how to circumvent a number of important theoretical and practical problems that we encountered. Using simulated data we show that model parameters seem consistent and appear unbiased in smaller data sets. We illustrate our methods using data from an ongoing cohort study.

Binomial Distribution↗