Development and validation of therapeutically relevant multi-gene biomarker classifiers.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Richard Simon.
Explore the source record for details and available documents.
MOTIVATION: In genomic studies, thousands of features are collected on relatively few samples. One of the goals of these studies is to build classifiers to predict the outcome of future observations. There are three inherent steps to this process: feature selection, model selection and prediction assessment. With a focus on prediction assessment, we compare several methods for estimating the 'true' prediction error of a prediction model in the presence of feature selection. RESULTS: For small studies where features are selected from thousands of candidates, the resubstitution and simple split-sample estimates are seriously biased. In these small samples, leave-one-out cross-validation (LOOCV), 10-fold cross-validation (CV) and the .632+ bootstrap have the smallest bias for diagonal discriminant analysis, nearest neighbor and classification trees. LOOCV and 10-fold CV have the smallest bias for linear discriminant analysis. Additionally, LOOCV, 5- and 10-fold CV, and the .632+ bootstrap have the lowest mean square error. The .632+ bootstrap is quite biased in small sample sizes with strong signal-to-noise ratios. Differences in performance among resampling methods are reduced as the number of specimens available increase. SUPPLEMENTARY INFORMATION: A complete compilation of results and R code for simulations and analyses are available in Molinaro et al. (2005) (http://linus.nci.nih.gov/brb/TechReport.htm).
PURPOSE: There is a wide spectrum of tumor responsiveness of rectal adenocarcinomas to preoperative chemoradiotherapy ranging from complete response to complete resistance. This study aimed to investigate whether parallel gene expression profiling of the primary tumor can contribute to stratification of patients into groups of responders or nonresponders. PATIENTS AND METHODS: Pretherapeutic biopsies from 30 locally advanced rectal carcinomas were analyzed for gene expression signatures using microarrays. All patients were participants of a phase III clinical trial (CAO/ARO/AIO-94, German Rectal Cancer Trial) and were randomized to receive a preoperative combined-modality therapy including fluorouracil and radiation. Class comparison was used to identify a set of genes that were differentially expressed between responders and nonresponders as measured by T level downsizing and histopathologic tumor regression grading. RESULTS: In an initial set of 23 patients, responders and nonresponders showed significantly different expression levels for 54 genes (P < .001). The ability to predict response to therapy using gene expression profiles was rigorously evaluated using leave-one-out cross-validation. Tumor behavior was correctly predicted in 83% of patients (P = .02). Sensitivity (correct prediction of response) was 78%, and specificity (correct prediction of nonresponse) was 86%, with a positive and negative predictive value of 78% and 86%, respectively. CONCLUSION: Our results suggest that pretherapeutic gene expression profiling may assist in response prediction of rectal adenocarcinomas to preoperative chemoradiotherapy. The implementation of gene expression profiles for treatment stratification and clinical management of cancer patients requires validation in large, independent studies, which are now warranted.
BACKGROUND: Normalization is a critical step in analysis of gene expression profiles. For dual-labeled arrays, global normalization assumes that the majority of the genes on the array are non-differentially expressed between the two channels and that the number of over-expressed genes approximately equals the number of under-expressed genes. These assumptions can be inappropriate for custom arrays or arrays in which the reference RNA is very different from the experimental samples. RESULTS: We propose a mixture model based normalization method that adaptively identifies non-differentially expressed genes and thereby substantially improves normalization for dual-labeled arrays in settings where the assumptions of global normalization are problematic. The new method is evaluated using both simulated and real data. CONCLUSIONS: The new normalization method is effective for general microarray platforms when samples with very different expression profile are co-hybridized and for custom arrays where the majority of genes are likely to be differentially expressed.
Therapeutic cancer vaccines have characteristics that require a new paradigm for phase I and phase II clinical development. Effective development plans may take advantage of some of the following observations: Dose ranging safety trials are not appropriate for many cancer vaccines. Dose ranging trials to establish an optimal biologic dose are often not practical. We have presented an efficient design of Korn et al. (4) to identify an immunogenic dose. Vaccine efficacy can be efficiently evaluated with tumor response as endpoint utilizing a two stage design with only 9 patients in the first stage. If no partial or complete responses are observed in the initial 9 patients, accrual to the trial is terminated. Optimization of vaccine delivery by comparing results of single arm phase II studies using immunological response as endpoint is problematic because of assay variation and potential non-comparability of patients in different studies. Randomized screening studies can be used to efficiently optimize vaccine immunogenicity. Efficiency in use of patients depends on having assay variation and inter-patient variability small relative to the difference in immunogenicity to be detected. Phase II studies using time to progression as endpoint are most interpretable if they employ randomized designs with a no-vaccine control group. Such designs may use an inflated type 1 error rate, and need not be prohibitively large if patients with rapidly progressive disease are studied. Interim monitoring plans may effectively limit the size of the trials by terminating accrual early when results are not consistent with the targeted improvement.
We used multistage models that incorporate the age dependent dynamics of normal breast tissue, clonal expansion of intermediate cells and mutational events to fit data for the age-specific incidence of breast cancers in the surveillance, epidemiology, and end results (SEER) registry. Our results suggest that two or three rate limiting events occurring at rates characteristic of point mutation rates for normal mammalian cells set in motion a sequence of other genomic changes that lead with high probability to breast carcinoma.
Explore the source record for details and available documents.
Determining sample sizes for microarray experiments is important but the complexity of these experiments, and the large amounts of data they produce, can make the sample size issue seem daunting, and tempt researchers to use rules of thumb in place of formal calculations based on the goals of the experiment. Here we present formulae for determining sample sizes to achieve a variety of experimental goals, including class comparison and the development of prognostic markers. Results are derived which describe the impact of pooling, technical replicates and dye-swap arrays on sample size requirements. These results are shown to depend on the relative sizes of different sources of variability. A variety of common types of experimental situations and designs used with single-label and dual-label microarrays are considered. We discuss procedures for controlling the false discovery rate. Our calculations are based on relatively simple yet realistic statistical models for the data, and provide straightforward sample size calculation formulae.
BACKGROUND: Patients with follicular lymphoma may survive for periods of less than 1 year to more than 20 years after diagnosis. We used gene-expression profiles of tumor-biopsy specimens obtained at diagnosis to develop a molecular predictor of the length of survival. METHODS: Gene-expression profiling was performed on 191 biopsy specimens obtained from patients with untreated follicular lymphoma. Supervised methods were used to discover expression patterns associated with the length of survival in a training set of 95 specimens. A molecular predictor of survival was constructed from these genes and validated in an independent test set of 96 specimens. RESULTS: Individual genes that predicted the length of survival were grouped into gene-expression signatures on the basis of their expression in the training set, and two such signatures were used to construct a survival predictor. The two signatures allowed patients with specimens in the test set to be divided into four quartiles with widely disparate median lengths of survival (13.6, 11.1, 10.8, and 3.9 years), independently of clinical prognostic variables. Flow cytometry showed that these signatures reflected gene expression by nonmalignant tumor-infiltrating immune cells. CONCLUSIONS: The length of survival among patients with follicular lymphoma correlates with the molecular features of nonmalignant immune cells present in the tumor at diagnosis.
Explore the source record for details and available documents.
PURPOSE: Genomic technologies make it increasingly possible to identify patients most likely to benefit from a molecularly targeted drug. This creates the opportunity to conduct targeted clinical trials with eligibility restricted to patients predicted to be responsive to the drug. EXPERIMENTAL DESIGN: We evaluated the relative efficiency of a targeted clinical trial design to an untargeted design for a randomized clinical trial comparing a new treatment to a control. Efficiency was evaluated with regard to number of patients required for randomization and number required for screening. RESULTS: The effectiveness of this design, relative to the more traditional design with broader eligibility, depends on multiple factors, including the proportion of responsive patients, the accuracy of the assay for predicting responsiveness, and the degree to which the mechanism of action of the drug is understood. Explicit formulas were derived for computing the relative efficiency of targeted versus untargeted designs. CONCLUSIONS: Targeted clinical trials can dramatically reduce the number of patients required for study in cases where the mechanism of action of the drug is understood and an accurate assay for responsiveness is available.
BACKGROUND: Clustering is one of the most commonly used methods for discovering hidden structure in microarray gene expression data. Most current methods for clustering samples are based on distance metrics utilizing all genes. This has the effect of obscuring clustering in samples that may be evident only when looking at a subset of genes, because noise from irrelevant genes dominates the signal from the relevant genes in the distance calculation. RESULTS: We describe an algorithm for automatically detecting clusters of samples that are discernable only in a subset of genes. We use iteration between Minimal Spanning Tree based clustering and feature selection to remove noise genes in a step-wise manner while simultaneously sharpening the clustering. Evaluation of this algorithm on synthetic data shows that it resolves planted clusters with high accuracy in spite of noise and the presence of other clusters. It also shows a low probability of detecting spurious clusters. Testing the algorithm on some well known micro-array data-sets reveals known biological classes as well as novel clusters. CONCLUSIONS: The iterative clustering method offers considerable improvement over clustering in all genes. This method can be used to discover partitions and their biological significance can be determined by comparing with clinical correlates and gene annotations. The MATLAB programs for the iterative clustering algorithm are available from http://linus.nci.nih.gov/supplement.html
Multiple sclerosis (MS) is an autoimmune disease in which myelin-specific T cells are believed to play a crucial pathogenic role. Nevertheless, so far it has been extremely difficult to demonstrate differences in T cell reactivity to myelin Ag between MS patients and controls. We believe that by using unphysiologically high Ag concentrations previous studies have missed a highly relevant aspect of autoimmune responses, i.e., T cells recognizing Ag with high functional avidity. Therefore, we focused on the characterization of high-avidity myelin-specific CD4+ T cells in a large cohort of MS patients and controls that was matched demographically and with respect to expression of MHC class II alleles. We demonstrated that their frequency is significantly higher in MS patients while the numbers of control T cells specific for influenza hemagglutinin are virtually identical between the two cohorts; that high-avidity T cells are enriched for previously in vivo-activated cells and are significantly skewed toward a proinflammatory phenotype. Moreover, the immunodominant epitopes that were most discriminatory between MS patients and controls differed from those described previously and were clearly biased toward epitopes with lower predicted binding affinities to HLA-DR molecules, pointing at the importance of thymic selection for the generation of the autoimmune T cell repertoire. Correlations between selected immunological parameters and magnetic resonance imaging markers indicate that the specificity and function of these cells influences phenotypic disease expression. These data have important implications for autoimmunity research and should be considered in the development of Ag-specific therapies in MS.
Studies on the elucidation of the specificity of the T cell receptor (TCR) at the antigen and peptide level have contributed to the current understanding of T cell cross-reactivity. Historically, most studies of T cell specificity and degeneracy have relied on the determination of the effects on T cell recognition of amino acid changes at individual positions or MHC binding residues, and thus they have been limited to a small set of possible ligands. Synthetic combinatorial libraries (SCLs), and in particular positional scanning synthetic combinatorial libraries (PS-SCLs) represent collections of millions to trillions of peptides which allow the unbiased elucidation of T cell ligands that stimulate clones of both known and unknown specificity. PS-SCLs have been used successfully to study T cell recognition and to identify and optimize T cell clone (TCC) epitopes in infectious diseases, autoimmune disorders and tumor antigens. PS-SCL-based biometrical analysis represents a further refinement in the analysis of the data derived from the screening of a library with a TCC. It combines this data with information derived from protein sequence databases to identify natural peptide ligands. PS-SCL-based biometrical analysis provides a method for the determination of new microbial antigen and autoantigen sequences based solely on functional data rather than sequence homology or motifs, making the method ideally suited for the prediction and identification of both native and cross-reactive epitopes by virtue of its ability to integrate the examination of trillions of peptides in a systematic manner with all of the protein sequences in a given database. We review here the application of PS-SCLs and biometrical analysis to identify cross-reactive T cell epitopes, as well as the current efforts to refine this strategy.
Peptides derived from pathogens or tumors are selectively presented by the major histocompatibility complex proteins (MHC) to the T lymphocytes. Antigenic peptide-MHC complexes on the cell surface are specifically recognized by T cells and, in conjunction with co-factor interactions, can activate the T cells to initiate the necessary immune response against the target cells. Peptides that are capable of binding to multiple MHC molecules are potential T cell epitopes for diverse human populations that may be useful in vaccine design. Bioinformatical approaches to predict MHC binding peptides can facilitate the resource-consuming effort of T cell epitope identification. We describe a new method for predicting MHC binding based on peptide property models constructed using biophysical parameters of the constituent amino acids and a training set of known binders. The models can be applied to development of anti-tumor vaccines by scanning proteins over-expressed in cancer cells for peptides that bind to a variety of MHC molecules. The complete algorithm is described and illustrated in the context of identifying candidate T cell epitopes for melanomas and breast cancers. We analyzed MART-1, S-100, MBP, and CD63 for melanoma and p53, MUC1, cyclin B1, HER-2/neu, and CEA for breast cancer. In general, proteins over-expressed in cancer cells may be identified using DNA microarray expression profiling. Comparisons of model predictions with available experimental data were assessed. The candidate epitopes identified by such a computational approach must be evaluated experimentally but the approach can provide an efficient and focused strategy for anti-cancer immunotherapy development.
Explore the source record for details and available documents.
We propose a new method for predicting MHC binding of peptides using biophysical parameters of the constituent amino acids. Unlike conventional matrix-based methods, our method does not assume independent binding of the individual side chains and uses a model that simultaneously represents all the residues. The model discovers the quantified 9-mer "property model" within the longer peptides that are most common among binders. Prediction for a new peptide is based on its statistical "distance" from the extracted peptide property model. MHC-specific peptide property models were constructed from compiled binder/nonbinder data using this method. We report the results of cross-validation of the prediction method and comparison with other methods. The comparison suggests that our method performs substantially better for some MHC class II molecules and equally well for other MHC types. To demonstrate large-scale utility, 30 HIV-1 reference genomes covering diverse subtypes were analyzed. Regions that are likely to bind MHC (A2, DR1, or DR4) and that are conserved across the HIV-1 subtypes were identified. These "epitope profiles" of the diverse HIV-1 strains can also be visually presented to facilitate discovery of conserved patterns naturally occurring in the viral genomes. As an essential step in designing vaccines, the revealed patterns may provide valuable information in identifying the immunologically important regions.
NF-kappaB is a transcription factor family that activates numerous genes that are related to cell survival, apoptosis, and cell migration. Its persistent activity is associated with tumor formation, growth, metastasis, and drug resistance in many cancer types, including lymphoma, colon cancer, and breast cancer. Current therapeutic efforts for inhibiting this central "switch" include using small molecules to block a selected target in this pathway. Recognizing the regulatory network structure of the NF-kappaB pathway, we examine in silico the effects of inhibitors targeting various network components, using a kinetic model of the pathway. By simulating the corresponding perturbed system dynamics, we show the resulting time course of inhibition has distinct target-specific profiles. In particular, greater oscillatory potential exists for inhibition of upstream events than for direct inhibition of NF-kappaB, at low drug concentrations. This phenomenon is observed also when we examine the dynamic effects of the recently approved proteasome inhibitor, bortezomib (PS-341), and compare it with other inhibitors, taking its pharmacokinetics into consideration. Such kinetic analyses of the "drugged" molecular system will facilitate optimal drug target selection and the development of treatment protocols for a molecularly targeted therapy.