Search PubMed⌕ Search

PubMed · 16504060

Tests for differential gene expression using weights in oligonucleotide microarray experiments.

Abstract

BACKGROUND: Microarray data analysts commonly filter out genes based on a number of ad hoc criteria prior to any high-level statistical analysis. Such ad hoc approaches could lead to conflicting conclusions with no clear guidance as to which method is most likely to be reproducible. Furthermore, the number of tests performed with concomitant inflation in type I error also plagues the statistical analysis of microarray data, since the number of tested quantities in a study significantly affects the family-wise error rate. It would, therefore, be very useful to develop and adopt strategies that allow quantification of the quality of each probeset, to filter out or give little credence to low-quality or unexpressed probesets, and to incorporate these strategies into gene selection within a multiple testing framework. RESULTS: We have proposed a unified scheme for filtering and gene selection. For Affymetrix gene expression microarrays, we developed new methods for measuring the reliability of a particular probeset in a single array, and we used these to develop measures for a set of arrays. These measures are then used as weights in standard t-statistic calculations, and are incorporated into the multiple testing procedures. We demonstrated the advantages of our methods using simulated data, publicly available spiked-in data as well as data comparing normal muscle to muscle from patients with Duchenne muscular dystrophy (DMD), in which a set of truly differentially expressed genes is known. CONCLUSION: Our quality measures provide convenient ways to search for individual genes of high quality. The quality weighting strategies we proposed for testing differential gene expression have demonstrable improvement on the traditional filtering methods, the standard t-statistic and a regularized t-statistic in Affymetrix data analysis.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pingzhao Hu, Joseph Beyene, Celia M T Greenwood. 2006-02-22. Tests for differential gene expression using weights in oligonucleotide microarray experiments.. https://doi.org/10.1186/1471-2164-7-33

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A nonlinear latent class model for joint analysis of multivariate longitudinal data and a binary outcome.

We consider a joint model for exploring association between several correlated longitudinal markers and a clinical event. A nonlinear growth mixture model exhibits the different latent classes of evolution of the latent quantity underlying the correlated longitudinal markers and a logistic regression models the probability of occurence of the clinical event according to the latent classes. By introducing a flexible nonlinear transformation including parameters to be estimated between each marker and the latent process, the model also deals with non-Gaussian continuous markers. Through an application on cognitive ageing, the two advantages of the model are underlined: (1) the latent profiles of evolution associated with the clinical event are described including covariate effects in the longitudinal model but also in the probability of class membership and in the probability of occurence of the event, and (2) a diagnostic and a prognostic tools are derived from the model for early detection of the clinical event using any available information about the longitudinal markers.

Data Interpretation, Statistical↗

Measurement of inter-rater agreement for transient events using Monte Carlo sampled permutations.

In this paper we demonstrate the adverse effect of serially observed data sequences containing transient events on the calculation of Cohen's kappa as an index of inter-rater agreement in the detection of these events. We develop and use a Monte-Carlo-based permutation technique to produce an empiric distribution of kappa in the presence of serial dependence. We find that the empiric confidence intervals for kappa tend to be wider than parametrically derived intervals and in the case of longer event lengths, are markedly so. We evaluate the effect of number and length of events, and further, describe and evaluate three permutation methods which match specific rating situations. Finally, we apply these techniques to the measurement of inter-rater agreement for sleep disordered breathing events, a transient event identified during nocturnal polysomnography, for which traditionally computed confidence intervals for kappa are incorrect.

Data Interpretation, Statistical↗