Search PubMedSearch

PubMed · 10474134

Some practical issues in binary data analysis.

Abstract

Three topics motivated by practical problems where the response variable is binary are described and illustrated. When a number of different explanatory variables are measured on each individual, a parsimonious model may be needed to predict the response of a future patient, or in selecting the variables that any treatment effect must be adjusted for. Some variable selection procedures used in conjunction with fitting logistic regression models are summarized and their performance investigated using a simulation study. A study to compare two devices for delivering anaesthetic gas to patients during surgery is then described, in which the response variable is the incidence of post-operative sore throat. In this study, the allocation of patient to device was non-random and a method for analysing these data that takes account of this aspect of the data is illustrated. In studies to compare different forms of contraceptive, the extent of regularity in the menstrual bleeding cycle is an important consideration for the acceptability of a contraceptive. Diary data on the menstrual bleeding pattern are therefore routinely collected. A method of summarizing the cyclic behaviour in the diary data for a particular woman is described, and extended to allow comparisons to be made between groups of women on different types of contraceptive. The approach is illustrated using a database made available by the World Health Organization.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

D Collett, K Stepniewska. Some practical issues in binary data analysis.. https://doi.org/10.1002/(sici)1097-0258(19990915%2F30)18%3A17%2F18%3C2209%3A%3Aaid-sim250%3E3.0.co%3B2-u

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation

Addressing current challenges in cancer immunotherapy with mathematical and computational modelling.

The goal of cancer immunotherapy is to boost a patient's immune response to a tumour. Yet, the design of an effective immunotherapy is complicated by various factors, including a potentially immunosuppressive tumour microenvironment, immune-modulating effects of conventional treatments and therapy-related toxicities. These complexities can be incorporated into mathematical and computational models of cancer immunotherapy that can then be used to aid in rational therapy design. In this review, we survey modelling approaches under the umbrella of the major challenges facing immunotherapy development, which encompass tumour classification, optimal treatment scheduling and combination therapy design. Although overlapping, each challenge has presented unique opportunities for modellers to make contributions using analytical and numerical analysis of model outcomes, as well as optimization algorithms. We discuss several examples of models that have grown in complexity as more biological information has become available, showcasing how model development is a dynamic process interlinked with the rapid advances in tumour-immune biology. We conclude the review with recommendations for modellers both with respect to methodology and biological direction that might help keep modellers at the forefront of cancer immunotherapy development.

Computer Simulation

Computer simulations of protein folding by targeted molecular dynamics.

We have performed 128 folding and 45 unfolding molecular dynamics runs of chymotrypsin inhibitor 2 (CI2) with an implicit solvation model for a total simulation time of 0.4 microseconds. Folding requires that the three-dimensional structure of the native state is known. It was simulated at 300 K by supplementing the force field with a harmonic restraint which acts on the root-mean-square deviation and allows to decrease the distance to the target conformation. High temperature and/or the harmonic restraint were used to induce unfolding. Of the 62 folding simulations started from random conformations, 31 reached the native structure, while the success rate was 83% for the 66 trajectories which began from conformations unfolded by high-temperature dynamics. A funnel-like energy landscape is observed for unfolding at 475 K, while the unfolding runs at 300 K and 375 K as well as most of the folding trajectories have an almost flat energy landscape for conformations with less than about 50% of native contacts formed. The sequence of events, i.e., secondary and tertiary structure formation, is similar in all folding and unfolding simulations, despite the diversity of the pathways. Previous unfolding simulations of CI2 performed with different force fields showed a similar sequence of events. These results suggest that the topology of the native state plays an important role in the folding process.

Computer Simulation