Search PubMed⌕ Search

Biomedical subjects

David M Rocke

Publications and source records attributed to David M Rocke.

7 recordsLinked to original sources

Estimation of transformation parameters for microarray data.

MOTIVATION AND RESULTS: Durbin et al. (2002), Huber et al. (2002) and Munson (2001) independently introduced a family of transformations (the generalized-log family) which stabilizes the variance of microarray data up to the first order. We introduce a method for estimating the transformation parameter in tandem with a linear model based on the procedure outlined in Box and Cox (1964). We also discuss means of finding transformations within the generalized-log family which are optimal under other criteria, such as minimum residual skewness and minimum mean-variance dependency. AVAILABILITY: R and Matlab code and test data are available from the authors on request.

Algorithms↗

Approximate variance-stabilizing transformations for gene-expression microarray data.

MOTIVATION: A variance stabilizing transformation for microarray data was recently introduced independently by several research groups. This transformation has sometimes been called the generalized logarithm or glog transformation. In this paper, we derive several alternative approximate variance stabilizing transformations that may be easier to use in some applications. RESULTS: We demonstrate that the started-log and the log-linear-hybrid transformation families can produce approximate variance stabilizing transformations for microarray data that are nearly as good as the generalized logarithm (glog) transformation. These transformations may be more convenient in some applications.

Algorithms↗

Improved significance test for DNA microarray data: temporal effects of shear stress on endothelial genes.

Statistical methods for identifying differentially expressed genes from microarray data are evolving. We developed a test for the statistical significance of differential expression as a function of time. When applied to microarray data obtained from endothelial cells exposed to shearing for different durations, the new multi-group test (G-test) identified three times as many genes as the one-way ANOVA at the same significance level. Using simulated data, we showed that this increase in sensitivity was achieved without sacrificing specificity. Several genes known to respond to shear stress by Northern blotting were identified by the G-test at P < or = 0.01 (but not by ANOVA), with similar temporal patterns. The validity and utility of the G-test were further supported by the examination of a few more example genes in relation to the present knowledge of their regulatory mechanisms. This new significance test may have broad application for the analysis of gene-expression studies and, in fact, to other biological studies in general.

Cells, Cultured↗

Multivariate survival analysis with doubly-censored data: application to the assessment of Accutane treatment for fibrodysplasia ossificans progressiva.

Fibrodysplasia ossificans progressiva is a rare genetic disorder in which the joints of patients become disabled by the formation of heterotopic bone. Data are available on the status of 11 joints of each of 21 patients before, during and after treatment with Accutane. These are compared with data obtained by questionnaire from 40 untreated patients to determine the efficacy of the treatment. Both left- and right-censoring are present in each group, which, together with the multivariate nature of the data and the time-dependent treatment covariate, makes analysis difficult. We consider two alternative parametric models for incorporating within-subject dependence: a marginal model and a frailty model. Both analyses suggest that Accutane treatment is effective. We discuss and illustrate the differences between the two approaches. We also discuss the extent to which the conclusions are compromised by the observational nature of the study.

Humans↗

Tumor classification by partial least squares using microarray gene expression data.

MOTIVATION: One important application of gene expression microarray data is classification of samples into categories, such as the type of tumor. The use of microarrays allows simultaneous monitoring of thousands of genes expressions per sample. This ability to measure gene expression en masse has resulted in data with the number of variables p(genes) far exceeding the number of samples N. Standard statistical methodologies in classification and prediction do not work well or even at all when N < p. Modification of existing statistical methodologies or development of new methodologies is needed for the analysis of microarray data. RESULTS: We propose a novel analysis procedure for classifying (predicting) human tumor samples based on microarray gene expressions. This procedure involves dimension reduction using Partial Least Squares (PLS) and classification using Logistic Discrimination (LD) and Quadratic Discriminant Analysis (QDA). We compare PLS to the well known dimension reduction method of Principal Components Analysis (PCA). Under many circumstances PLS proves superior; we illustrate a condition when PCA particularly fails to predict well relative to PLS. The proposed methods were applied to five different microarray data sets involving various human tumor samples: (1) normal versus ovarian tumor; (2) Acute Myeloid Leukemia (AML) versus Acute Lymphoblastic Leukemia (ALL); (3) Diffuse Large B-cell Lymphoma (DLBCLL) versus B-cell Chronic Lymphocytic Leukemia (BCLL); (4) normal versus colon tumor; and (5) Non-Small-Cell-Lung-Carcinoma (NSCLC) versus renal samples. Stability of classification results and methods were further assessed by re-randomization studies.

Colonic Neoplasms↗

Partial least squares proportional hazard regression for application to DNA microarray survival data.

MOTIVATION: Microarrays are increasingly used in cancer research. When gene transcription data from microarray experiments also contains patient survival information, it is often of interest to predict the survival times based on the gene expression. In this paper we consider the well-known proportional hazard (PH) regression model for survival analysis. Ordinarily, the PH model is used with a few covariates and many observations (subjects). We consider here the case that the number of covariates, p, exceeds the number of samples, N, a setting typical of gene expression data from DNA microarrays. RESULTS: For a given vector of response values which are survival times and p gene expressions (covariates) we examine the problem of how to predict the survival probabilities, when N << p. The approach taken to cope with the high dimensionality is to reduce the dimension using partial least squares with the response variable as the vector of survival times. After dimension reduction, the extracted PLS gene components are then used as covariates in a PH regression to predict the survival probabilities. We demonstrate the use of the methodology on two cDNA gene expression data sets, both containing survival data. The first data set contains 40 diffuse large B-cell lymphoma (DLBCL) tissue samples and the second data set contains 49 tissue samples from patients with locally advanced breast cancer in a prospective study.

Breast Neoplasms↗

Multi-class cancer classification via partial least squares with gene expression profiles.

MOTIVATION: Discrimination between two classes such as normal and cancer samples and between two types of cancers based on gene expression profiles is an important problem which has practical implications as well as the potential to further our understanding of gene expression of various cancer cells. Classification or discrimination of more than two groups or classes (multi-class) is also needed. The need for multi-class discrimination methodologies is apparent in many microarray experiments where various cancer types are considered simultaneously. RESULTS: Thus, in this paper we present the extension to the classification methodology proposed earlier Nguyen and Rocke (2002b; Bioinformatics, 18, 39-50) to classify cancer samples from multiple classes. The methodologies proposed in this paper are applied to four gene expression data sets with multiple classes: (a) a hereditary breast cancer data set with (1) BRCA1-mutation, (2) BRCA2-mutation and (3) sporadic breast cancer samples, (b) an acute leukemia data set with (1) acute myeloid leukemia (AML), (2) T-cell acute lymphoblastic leukemia (T-ALL) and (3) B-cell acute lymphoblastic leukemia (B-ALL) samples, (c) a lymphoma data set with (1) diffuse large B-cell lymphoma (DLBCL), (2) B-cell chronic lymphocytic leukemia (BCLL) and (3) follicular lymphoma (FL) samples, and (d) the NCI60 data set with cell lines derived from cancers of various sites of origin. In addition, we evaluated the classification algorithms and examined the variability of the error rates using simulations based on randomization of the real data sets. We note that there are other methods for addressing multi-class prediction recently and our approach is along the line of Nguyen and Rocke (2002b; Bioinformatics, 18, 39-50). CONTACT: dnguyen@stat.tamu.edu; dmrocke@ucdavis.edu

Algorithms↗