Search PubMed⌕ Search

Biomedical subjects

G A Whitmore

Publications and source records attributed to G A Whitmore.

12 recordsLinked to original sources

Power and sample size for DNA microarray studies.

A microarray study aims at having a high probability of declaring genes to be differentially expressed if they are truly expressed, while keeping the probability of making false declarations of expression acceptably low. Thus, in formal terms, well-designed microarray studies will have high power while controlling type I error risk. Achieving this objective is the purpose of this paper. Here, we discuss conceptual issues and present computational methods for statistical power and sample size in microarray studies, taking account of the multiple testing that is generic to these studies. The discussion encompasses choices of experimental design and replication for a study. Practical examples are used to demonstrate the methods. The examples show forcefully that replication of a microarray experiment can yield large increases in statistical power. The paper refers to cDNA arrays in the discussion and illustrations but the proposed methodology is equally applicable to expression data from oligonucleotide arrays.

Analysis of Variance↗

Models for microarray gene expression data.

This paper describes a general methodology for the analysis of differential gene expression based on microarray data. First, we characterize the data by a linear statistical model that accounts for relevant sources of variation in the data and then we consider estimation of the model parameters. Because microarray studies typically involve thousands of genes, we propose a two-stage method for parameter estimation. The interaction terms for genes and experimental conditions in this model capture all relevant information about differential gene expression in the microarray data. We propose a mixture distribution model for a summary statistic of differential expression that consists of null and alternative component distributions. The mixture model suggests two methods for identifying genes exhibiting differential expression. One is a frequentist method that identifies distinguished genes and the other an empirical Bayes procedure that yields estimated posterior probabilities of differential expression, conditional on observed microarray readings.

Animals↗

A statistical model for investigating binding probabilities of DNA nucleotide sequences using microarrays.

There is considerable scientific interest in knowing the probability that a site-specific transcription factor will bind to a given DNA sequence. Microarray methods provide an effective means for assessing the binding affinities of a large number of DNA sequences as demonstrated by Bulyk et al. (2001, Proceedings of the National Academy of Sciences, USA 98, 7158-7163) in their study of the DNA-binding specificities of Zif268 zinc fingers using microarray technology. In a follow-up investigation, Bulyk, Johnson, and Church (2002, Nucleic Acid Research 30, 1255-1261) studied the interdependence of nucleotides on the binding affinities of transcription proteins. Our article is motivated by this pair of studies. We present a general statistical methodology for analyzing microarray intensity measurements reflecting DNA-protein interactions. The log probability of a protein binding to a DNA sequence on an array is modeled using a linear ANOVA model. This model is convenient because it employs familiar statistical concepts and procedures and also because it is effective for investigating the probability structure of the binding mechanism.

Analysis of Variance↗

Importance of replication in microarray gene expression studies: statistical methods and evidence from repetitive cDNA hybridizations.

We present statistical methods for analyzing replicated cDNA microarray expression data and report the results of a controlled experiment. The study was conducted to investigate inherent variability in gene expression data and the extent to which replication in an experiment produces more consistent and reliable findings. We introduce a statistical model to describe the probability that mRNA is contained in the target sample tissue, converted to probe, and ultimately detected on the slide. We also introduce a method to analyze the combined data from all replicates. Of the 288 genes considered in this controlled experiment, 32 would be expected to produce strong hybridization signals because of the known presence of repetitive sequences within them. Results based on individual replicates, however, show that there are 55, 36, and 58 highly expressed genes in replicates 1, 2, and 3, respectively. On the other hand, an analysis by using the combined data from all 3 replicates reveals that only 2 of the 288 genes are incorrectly classified as expressed. Our experiment shows that any single microarray output is subject to substantial variability. By pooling data from replicates, we can provide a more reliable analysis of gene expression data. Therefore, we conclude that designing experiments with replications will greatly reduce misclassification rates. We recommend that at least three replicates be used in designing experiments by using cDNA microarrays, particularly when gene expression data from single specimens are being analyzed.

DNA, Complementary↗

Statistical inference for serial dilution assay data.

Serial dilution assays are widely employed for estimating substance concentrations and minimum inhibitory concentrations. The Poisson-Bernoulli model for such assays is appropriate for count data but not for continuous measurements that are encountered in applications involving substance concentrations. This paper presents practical inference methods based on a log-normal model and illustrates these methods using a case application involving bacterial toxins.

Algorithms↗

Failure inference from a marker process based on a bivariate Wiener model.

Many models have been proposed that relate failure times and stochastic time-varying covariates. In some of these models, failure occurs when a particular observable marker crosses a threshold level. We are interested in the more difficult, and often more realistic, situation where failure is not related deterministically to an observable marker. In this case, joint models for marker evolution and failure tend to lead to complicated calculations for characteristics such as the marginal distribution of failure time or the joint distribution of failure time and marker value at failure. This paper presents a model based on a bivariate Wiener process in which one component represents the marker and the second, which is latent (unobservable), determines the failure time. In particular, failure occurs when the latent component crosses a threshold level. The model yields reasonably simple expressions for the characteristics mentioned above and is easy to fit to commonly occurring data that involve the marker value at the censoring time for surviving cases and the marker value and failure time for failing cases. Parametric and predictive inference are discussed, as well as model checking. An extension of the model permits the construction of a composite marker from several candidate markers that may be available. The methodology is demonstrated by a simulated example and a case application.

Biometry↗

Modelling accelerated degradation data using Wiener diffusion with a time scale transformation.

Engineering degradation tests allow industry to assess the potential life span of long-life products that do not fail readily under accelerated conditions in life tests. A general statistical model is presented here for performance degradation of an item of equipment. The degradation process in the model is taken to be a Wiener diffusion process with a time scale transformation. The model incorporates Arrhenius extrapolation for high stress testing. The lifetime of an item is defined as the time until performance deteriorates to a specified failure threshold. The model can be used to predict the lifetime of an item or the extent of degradation of an item at a specified future time. Inference methods for the model parameters, based on accelerated degradation test data, are presented. The model and inference methods are illustrated with a case application involving self-regulating heating cables. The paper also discusses a number of practical issues encountered in applications.

Equipment Failure↗

Analysis of overdispersed count data by mixtures of Poisson variables and Poisson processes.

Count data often show overdispersion compared to the Poisson distribution. Overdispersion is typically modeled by a random effect for the mean, based on the gamma distribution, leading to the negative binomial distribution for the count. This paper considers a larger family of mixture distributions, including the inverse Gaussian mixture distribution. It is demonstrated that it gives a significantly better fit for a data set on the frequency of epileptic seizures. The same approach can be used to generate counting processes from Poisson processes, where the rate or the time is random. A random rate corresponds to variation between patients, whereas a random time corresponds to variation within patients.

Anticonvulsants↗

Estimating degradation by a Wiener diffusion process subject to measurement error.

Most materials and components degrade physically before they fail. Engineering degradation tests are designed to measure these degradation processes. Measurements in the tests reflect the inherent randomness of degradation itself as well as measurement errors created by imperfect instruments, procedures and environments. This paper describes a statistical model for measured degradation data that takes both sources of variation into account. The degradation process in the model is taken to be a Wiener diffusion process. The measurement errors are assumed to be independent normal random outcomes that are independent of the degradation process. The paper describes inference procedures for the model and discusses some practical issues that must be considered in dealing with the statistical problem. A case study is presented.

Engineering↗

The mortality component of health status indexes.

The mortality component of contemporary health indexes is discussed. Since these indexes reduce to mortality indexes when only life and death states enter the analysis, they share the conceptual weaknesses of mortality indexes. Also, they do not incorporate consumption variables explicity and therefore provide no structure for relating health status and living standard. Some attention is devoted to methodological problems of assessing survival probabilities, either from survey or experimental data or from beliefs of experts or individuals who are affected directly. The final section deals with individual preferences for survival lotteries. Conceptual weaknesses of common indexes are discussed, several canonical models for survival preferences are presented, the interdependence of individual utilities is discussed, and methods for eliciting individual survival preferences are considered, along with some illustrative empirical results.

Choice Behavior↗

The inverse Gaussian distribution as a model of hospital stay.

Properties of the inverse gaussian distribution are presented with comments on fitting the distribution to lentgh-of-stay data. A conceptual framework for the hospitalization process is described; it suggests that the inverse gaussian distribution has considerable potential as both a descriptive and prescriptive model of length of stay, especially in the setting of psychiatric hospitals.

Humans↗