Search PubMed⌕ Search

PubMed · 16802876

Some statistical issues in microarray gene expression data.

Abstract

In this paper we discuss some of the statistical issues that should be considered when conducting experiments involving microarray gene expression data. We discuss statistical issues related to preprocessing the data as well as the analysis of the data. Analysis of the data is discussed in three contexts: class comparison, class prediction and class discovery. We also review the methods used in two studies that are using microarray gene expression to assess the effect of exposure to radiofrequency (RF) fields on gene expression. Our intent is to provide a guide for radiation researchers when conducting studies involving microarray gene expression data.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Matthew S Mayo, Byron J Gajewski, Jeffrey S Morris. 2006. Some statistical issues in microarray gene expression data.. https://doi.org/10.1667/rr3576.1

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Striping artifact removal in VisiumHD data through nuclear counts modeling.

MOTIVATION: 10x Genomics VisiumHD enables spatial transcriptomics at 2 µm × 2 µm resolution but exhibits slide-specific, non-periodic striping artifacts due to lane-width variability. These multiplicative row/column effects distort bin total counts and can bias downstream analyses. The state-of-the-art destriping approach is the normalization procedure used as a preprocessing step in bin2cell; it applies sequential high-quantile row- then column-wise normalization, which is asymmetric and can introduce edge effects/macro-stripes and distortions of large-scale total-count structure. RESULTS: We propose a statistical destriping approach that leverages nuclei segmentation from the co-registered H&E image. Assuming transcript abundance is constant within each nucleus, we model bin counts with a negative binomial distribution whose mean is a product of a nucleus-specific concentration and row- and column-specific stripe-factors reflecting lane-width variation. We fit all parameters in a generalized linear modeling framework with cross-validated regularization on stripe-factors and iterative dispersion estimation, and use the fitted parameters to correct the observed counts into a destriped image. On synthetic data with known ground truth, our method improves stripe-factor estimation accuracy and reduces error in corrected counts relative to bin2cell and bin2cell-derived baselines. Across four public VisiumHD slides, it consistently lowers striping intensity while substantially better preserving biological signal present in the large-scale global count structure and avoiding the artifacts introduced by other methods. AVAILABILITY AND IMPLEMENTATION: All source code and links to publicly available data used for this study are available at https://github.com/paolamalsot/destriping-GLM.

Artifacts↗

Cleavage of cystatin C is not associated with multiple sclerosis.

Recently, Irani and colleagues proposed a C-terminal cleaved isoform cystatin C (12.5 kDa) in cerebrospinal fluid as a marker of multiple sclerosis. In this study, we demonstrate that the 12.5 kDa product of cystatin C is formed by degradation of the first eight N-terminal residues. Moreover, such a degradation is not specific in the cerebrospinal fluid of multiple sclerosis, but rather is given by an inappropriate sample storage at -20 degrees C. We conclude that the use of the 12.5 kDa product of cystatin C in cerebrospinal fluid might lead to a fallacious diagnosis of multiple sclerosis. Preanalytical validation procedure is mandatory for proteomics investigations.

Artifacts↗