Search PubMed⌕ Search

Biomedical subjects

Richard M Simon

Publications and source records attributed to Richard M Simon.

8 recordsLinked to original sources

Critical review of published microarray studies for cancer outcome and guidelines on statistical analysis and reporting.

BACKGROUND: Both the validity and the reproducibility of microarray-based clinical research have been challenged. There is a need for critical review of the statistical analysis and reporting in published microarray studies that focus on cancer-related clinical outcomes. METHODS: Studies published through 2004 in which microarray-based gene expression profiles were analyzed for their relation to a clinical cancer outcome were identified through a Medline search followed by hand screening of abstracts and full text articles. Studies that were eligible for our analysis addressed one or more outcomes that were either an event occurring during follow-up, such as death or relapse, or a therapeutic response. We recorded descriptive characteristics for all the selected studies. A critical review of outcome-related statistical analyses was undertaken for the articles published in 2004. RESULTS: Ninety studies were identified, and their descriptive characteristics are presented. Sixty-eight (76%) were published in journals of impact factor greater than 6. A detailed account of the 42 studies (47%) published in 2004 is reported. Twenty-one (50%) of them contained at least one of the following three basic flaws: 1) in outcome-related gene finding, an unstated, unclear, or inadequate control for multiple testing; 2) in class discovery, a spurious claim of correlation between clusters and clinical outcome, made after clustering samples using a selection of outcome-related differentially expressed genes; or 3) in supervised prediction, a biased estimation of the prediction accuracy through an incorrect cross-validation procedure. CONCLUSIONS: The most common and serious mistakes and misunderstandings recorded in published studies are described and illustrated. Based on this analysis, a proposal of guidelines for statistical analysis and reporting for clinical microarray studies, presented as a checklist of "Do's and Don'ts," is provided.

Bias↗

Sample size planning for developing classifiers using high-dimensional DNA microarray data.

Many gene expression studies attempt to develop a predictor of pre-defined diagnostic or prognostic classes. If the classes are similar biologically, then the number of genes that are differentially expressed between the classes is likely to be small compared to the total number of genes measured. This motivates a two-step process for predictor development, a subset of differentially expressed genes is selected for use in the predictor and then the predictor constructed from these. Both these steps will introduce variability into the resulting classifier, so both must be incorporated in sample size estimation. We introduce a methodology for sample size determination for prediction in the context of high-dimensional data that captures variability in both steps of predictor development. The methodology is based on a parametric probability model, but permits sample size computations to be carried out in a practical manner without extensive requirements for preliminary data. We find that many prediction problems do not require a large training set of arrays for classifier development.

Computer Simulation↗

Gene expression patterns and profile changes pre- and post-erlotinib treatment in patients with metastatic breast cancer.

PURPOSE: To delineate gene expression patterns and profile changes in metastatic tumor biopsies at baseline and 1 month after treatment with the epidermal growth factor receptor (EGFR) tyrosine kinase inhibitor erlotinib in patients with metastatic breast cancer. EXPERIMENTAL DESIGN: Patients were treated with 150 mg of oral erlotinib daily. Gene expression profiles were measured with Affymetrix U133A GeneChip and immunohistochemistry was used to validate microarray findings. RESULTS: Estrogen receptor (ER) status by immunohistochemistry is nearly coincided with the two major expression clusters determined by expression of genes using unsupervised hierarchical clustering analysis. One of 10 patients had an EGFR-positive tumor detected by both microarray and immunohistochemistry. In this tumor, tissue inhibitor of metalloproteinases-3 and collagen type 1 alpha 2, which are the EGF-down-regulated growth repressors, were significantly increased by erlotinib. Gene changes in EGFR-negative tumors are those of G-protein-linked and cell surface receptor-linked signaling. Gene ontology comparison analysis pretreatment and posttreatment in EGFR-negative tumors revealed biological process categories that have more genes differentially expressed than expected by chance. Among 495 gene ontology categories, the significant differed gene ontology groups include G-protein-coupled receptor protein signaling (34 genes, P = 0.002) and cell surface receptor-linked signal transduction (74 genes, P = 0.007). CONCLUSIONS: ER status reflects the major difference in gene expression pattern in metastatic breast cancer. Erlotinib had effects on genes of EGFR signaling pathway in the EGFR-positive tumor and on gene ontology biological process categories or genes that have function in signal transduction in EGFR-negative tumors.

Biomarkers, Tumor↗

Interlaboratory comparability study of cancer gene expression analysis using oligonucleotide microarrays.

A key step in bringing gene expression data into clinical practice is the conduct of large studies to confirm preliminary models. The performance of such confirmatory studies and the transition to clinical practice requires that microarray data from different laboratories are comparable and reproducible. We designed a study to assess the comparability of data from four laboratories that will conduct a larger microarray profiling confirmation project in lung adenocarcinomas. To test the feasibility of combining data across laboratories, frozen tumor tissues, cell line pellets, and purified RNA samples were analyzed at each of the four laboratories. Samples of each type and several subsamples from each tumor and each cell line were blinded before being distributed. The laboratories followed a common protocol for all steps of tissue processing, RNA extraction, and microarray analysis using Affymetrix Human Genome U133A arrays. High within-laboratory and between-laboratory correlations were observed on the purified RNA samples, the cell lines, and the frozen tumor tissues. Intraclass correlation within laboratories was only slightly stronger than between laboratories, and the intraclass correlation tended to be weakest for genes expressed at low levels and showing small variation. Finally, hierarchical cluster analysis revealed that the repeated samples clustered together regardless of the laboratory in which the experiments were done. The findings indicate that under properly controlled conditions it is feasible to perform complete tumor microarray analysis, from tissue processing to hybridization and scanning, at multiple independent laboratories for a single study.

Adenocarcinoma↗

A random variance model for detection of differential gene expression in small microarray experiments.

MOTIVATION: Microarray techniques provide a valuable way of characterizing the molecular nature of disease. Unfortunately expense and limited specimen availability often lead to studies with small sample sizes. This makes accurate estimation of variability difficult, since variance estimates made on a gene by gene basis will have few degrees of freedom, and the assumption that all genes share equal variance is unlikely to be true. RESULTS: We propose a model by which the within gene variances are drawn from an inverse gamma distribution, whose parameters are estimated across all genes. This results in a test statistic that is a minor variation of those used in standard linear models. We demonstrate that the model assumptions are valid on experimental data, and that the model has more power than standard tests to pick up large changes in expression, while not increasing the rate of false positives. AVAILABILITY: This method is incorporated into BRB-ArrayTools version 3.0 (http://linus.nci.nih.gov/BRB-ArrayTools.html). SUPPLEMENTARY MATERIAL: ftp://linus.nci.nih.gov/pub/techreport/RVM_supplement.pdf

Algorithms↗