Search PubMed⌕ Search

Biomedical subjects

R Tibshirani

Publications and source records attributed to R Tibshirani.

At least 19 recordsLinked to original sources

Gene expression patterns of breast carcinomas distinguish tumor subclasses with clinical implications.

The purpose of this study was to classify breast carcinomas based on variations in gene expression patterns derived from cDNA microarrays and to correlate tumor characteristics to clinical outcome. A total of 85 cDNA microarray experiments representing 78 cancers, three fibroadenomas, and four normal breast tissues were analyzed by hierarchical clustering. As reported previously, the cancers could be classified into a basal epithelial-like group, an ERBB2-overexpressing group and a normal breast-like group based on variations in gene expression. A novel finding was that the previously characterized luminal epithelial/estrogen receptor-positive group could be divided into at least two subgroups, each with a distinctive expression profile. These subtypes proved to be reasonably robust by clustering using two different gene sets: first, a set of 456 cDNA clones previously selected to reflect intrinsic properties of the tumors and, second, a gene set that highly correlated with patient outcome. Survival analyses on a subcohort of patients with locally advanced breast cancer uniformly treated in a prospective study showed significantly different outcomes for the patients belonging to the various groups, including a poor prognosis for the basal-like subtype and a significant difference in outcome for the two estrogen receptor-positive groups.

Algorithms↗

Expression of a single gene, BCL-6, strongly predicts survival in patients with diffuse large B-cell lymphoma.

Diffuse large B-cell lymphoma (DLBCL) is characterized by a marked degree of morphologic and clinical heterogeneity. Establishment of parameters that can predict outcome could help to identify patients who may benefit from risk-adjusted therapies. BCL-6 is a proto-oncogene commonly implicated in DLBCL pathogenesis. A real-time reverse transcription-polymerase chain reaction assay was established for accurate and reproducible determination of BCL-6 mRNA expression. The method was applied to evaluate the prognostic significance of BCL-6 expression in DLBCL. BCL-6 mRNA expression was assessed in tumor specimens obtained at the time of diagnosis from 22 patients with primary DLBCL. All patients were subsequently treated with anthracycline-based chemotherapy regimens. These patients could be divided into 2 DLBCL subgroups, one with high BCL-6 gene expression whose median overall survival (OS) time was 171 months and the other with low BCL-6 gene expression whose median OS was 24 months (P =.007). BCL-6 gene expression also predicted OS in an independent validation set of 39 patients with primary DLBCL (P =.01). BCL-6 protein expression, assessed by immunohistochemistry, also predicted longer OS in patients with DLBCL. BCL-6 gene expression was an independent survival predicting factor in multivariate analysis together with the elements of the International Prognostic Index (IPI) (P =.038). By contrast, the aggregate IPI score did not add further prognostic information to the patients' stratification by BCL-6 gene expression. High BCL-6 mRNA expression should be considered a new favorable prognostic factor in DLBCL and should be used in the stratification and the design of risk-adjusted therapies for patients with DLBCL. (Blood. 2001;98:945-951)

Adolescent↗

Significance analysis of microarrays applied to the ionizing radiation response.

Microarrays can measure the expression of thousands of genes to identify changes in expression between different biological states. Methods are needed to determine the significance of these changes while accounting for the enormous number of genes. We describe a method, Significance Analysis of Microarrays (SAM), that assigns a score to each gene on the basis of change in gene expression relative to the standard deviation of repeated measurements. For genes with scores greater than an adjustable threshold, SAM uses permutations of the repeated measurements to estimate the percentage of genes identified by chance, the false discovery rate (FDR). When the transcriptional response of human cells to ionizing radiation was measured by microarrays, SAM identified 34 genes that changed at least 1.5-fold with an estimated FDR of 12%, compared with FDRs of 60 and 84% by using conventional methods of analysis. Of the 34 genes, 19 were involved in cell cycle regulation and 3 in apoptosis. Surprisingly, four nucleotide excision repair genes were induced, suggesting that this repair pathway for UV-damaged DNA might play a previously unrecognized role in repairing DNA damaged by ionizing radiation.

Apoptosis↗

Supervised harvesting of expression trees.

BACKGROUND: We propose a new method for supervised learning from gene expression data. We call it 'tree harvesting'. This technique starts with a hierarchical clustering of genes, then models the outcome variable as a sum of the average expression profiles of chosen clusters and their products. It can be applied to many different kinds of outcome measures such as censored survival times, or a response falling in two or more classes (for example, cancer classes). The method can discover genes that have strong effects on their own, and genes that interact with other genes. RESULTS: We illustrate the method on data from a lymphoma study, and on a dataset containing samples from eight different cancers. It identified some potentially interesting gene clusters. In simulation studies we found that the procedure may require a large number of experimental samples to successfully discover interactions. CONCLUSIONS: Tree harvesting is a potentially useful tool for exploration of gene expression data and identification of interesting clusters of genes worthy of further investigation.

Gene Expression Profiling↗

Missing value estimation methods for DNA microarrays.

MOTIVATION: Gene expression microarray experiments can generate data sets with multiple missing expression values. Unfortunately, many algorithms for gene expression analysis require a complete matrix of gene array values as input. For example, methods such as hierarchical clustering and K-means clustering are not robust to missing data, and may lose effectiveness even with a few missing values. Methods for imputing missing data are needed, therefore, to minimize the effect of incomplete data sets on analyses, and to increase the range of data sets to which these algorithms can be applied. In this report, we investigate automated methods for estimating missing data. RESULTS: We present a comparative study of several methods for the estimation of missing values in gene microarray data. We implemented and evaluated three methods: a Singular Value Decomposition (SVD) based method (SVDimpute), weighted K-nearest neighbors (KNNimpute), and row average. We evaluated the methods using a variety of parameter settings and over different real data sets, and assessed the robustness of the imputation methods to the amount of missing data over the range of 1--20% missing values. We show that KNNimpute appears to provide a more robust and sensitive method for missing value estimation than SVDimpute, and both SVDimpute and KNNimpute surpass the commonly used row average method (as well as filling missing values with zeros). We report results of the comparative experiments and provide recommendations and tools for accurate estimation of missing microarray data under a variety of conditions.

Algorithms↗

The inference of antigen selection on Ig genes.

Analysis of somatic mutations in V regions of Ig genes is important for understanding various biological processes. It is customary to estimate Ag selection on Ig genes by assessment of replacement (R) as opposed to silent (S) mutations in the complementary-determining regions and S as opposed to R mutations in the framework regions. In the past such an evaluation was performed using a binomial distribution model equation, which is inappropriate for Ig genes in which mutations have four different distribution possibilities (R and S mutations in the complementary-determining region and/or framework regions of the gene). In the present work, we propose a multinomial distribution model for assessment of Ag selection. Side-by-side application of multinomial and binomial models on 86 previously established Ig sequences disclosed 8 discrepancies, leading to opposite statistical conclusions about Ag selection. We suggest the use of the multinomial model for all future analysis of Ag selection.

Antigens↗

'Gene shaving' as a method for identifying distinct sets of genes with similar expression patterns.

BACKGROUND: Large gene expression studies, such as those conducted using DNA arrays, often provide millions of different pieces of data. To address the problem of analyzing such data, we describe a statistical method, which we have called 'gene shaving'. The method identifies subsets of genes with coherent expression patterns and large variation across conditions. Gene shaving differs from hierarchical clustering and other widely used methods for analyzing gene expression studies in that genes may belong to more than one cluster, and the clustering may be supervised by an outcome measure. The technique can be 'unsupervised', that is, the genes and samples are treated as unlabeled, or partially or fully supervised by using known properties of the genes or samples to assist in finding meaningful groupings. RESULTS: We illustrate the use of the gene shaving method to analyze gene expression measurements made on samples from patients with diffuse large B-cell lymphoma. The method identifies a small cluster of genes whose expression is highly predictive of survival. CONCLUSIONS: The gene shaving method is a potentially useful tool for exploration of gene expression data and identification of interesting clusters of genes worth further investigation.

Algorithms↗

Molecular analysis of immunoglobulin genes in diffuse large B-cell lymphomas.

Diffuse large B-cell lymphoma (DLBCL) is a common type of non-Hodgkin's lymphoma (NHL) that is highly heterogeneous from both clinical and histopathologic viewpoints. The immunoglobulin (Ig) heavy (H) chain variable region genes were examined in 71 patients with untreated primary DLBCL. Fifty-eight potentially functional V(H) genes were detected in 53 DLBCL cases; V(H) genes were nonfunctional in 9 cases and were not detected in an additional 9 cases. The use of V(H) gene families by DLBCL tumors was unbiased without overrepresentation of any particular V(H) gene or gene family. Analysis of Ig mutations in comparison to the most closely related germline gene disclosed mutated V(H) genes in all but 1 DLBCL case. More than 2% difference from the most similar germline sequence was detected in 52 potentially functional and the 8 nonfunctional V(H) gene sequences, whereas less than 2% difference from the germline sequence was observed in 3 V(H) gene isolates. Only 3 V(H) gene isolates were unmutated. No correlation was found between V(H) gene use, mutation level, and International Prognostic Index (IPI) or survival. Six of 8 tested tumors showed evidence of ongoing somatic mutations. Evidence for positive or negative antigen selection pressure was observed in 65% of mutated DLBCL cases. Our findings indicate that the etiology and the driving forces for clonal expansion are heterogeneous, which may explain the well-known clinical and pathologic heterogeneity of DLBCL. (Blood. 2000;95:1797-1803)

B-Lymphocytes↗

Distinct types of diffuse large B-cell lymphoma identified by gene expression profiling.

Diffuse large B-cell lymphoma (DLBCL), the most common subtype of non-Hodgkin's lymphoma, is clinically heterogeneous: 40% of patients respond well to current therapy and have prolonged survival, whereas the remainder succumb to the disease. We proposed that this variability in natural history reflects unrecognized molecular heterogeneity in the tumours. Using DNA microarrays, we have conducted a systematic characterization of gene expression in B-cell malignancies. Here we show that there is diversity in gene expression among the tumours of DLBCL patients, apparently reflecting the variation in tumour proliferation rate, host response and differentiation state of the tumour. We identified two molecularly distinct forms of DLBCL which had gene expression patterns indicative of different stages of B-cell differentiation. One type expressed genes characteristic of germinal centre B cells ('germinal centre B-like DLBCL'); the second type expressed genes normally induced during in vitro activation of peripheral blood B cells ('activated B-like DLBCL'). Patients with germinal centre B-like DLBCL had a significantly better overall survival than those with activated B-like DLBCL. The molecular classification of tumours on the basis of gene expression can thus identify previously undetected and clinically significant subtypes of cancer.

Adult↗

Childhood leukemia and personal monitoring of residential exposures to electric and magnetic fields in Ontario, Canada.

OBJECTIVES: To evaluate the risk of childhood leukemia in relation to residential electric and magnetic field (EMF) exposures. METHODS: A case control study based on 88 cases and 133 controls used different assessment methods to determine EMF exposure in the child's current residence. Cases comprised incident leukemias diagnosed at 0-14 years of age between 1985-1993 from a larger study in southern Ontario; population controls were individually matched to the cases by age and sex. Exposure was measured by a personal monitoring device worn by the child during usual activities at home, by point-in-time measurements in three rooms and according to wire code assigned to the child's residence. RESULTS: An association between magnetic field exposures as measured with the personal monitor and increased risk of leukemia was observed. The risk was more pronounced for those children diagnosed at less than 6 years of age and those with acute lymphoblastic leukemia. Risk estimates associated with magnetic fields tended to increase after adjusting for power consumption and potential confounders with significant odds ratios (OR) (OR: 4.5, 95% confidence interval (CI): 1.3-15.9) observed for exposures > or = 0.14 microTesla (microT). For the most part point-in-time measurements of magnetic fields were associated with non-significant elevations in risk which were generally compatible with previous research. Residential proximity to power lines having a high current configuration was not associated with increased risk of leukemia. Exposures to electric fields as measured by personal monitoring were associated with a decreased leukemia risk. CONCLUSIONS: The findings relating to magnetic field exposures directly measured by personal monitoring support an association with the risk of childhood leukemia. As exposure assessment is refined, the possible role of magnetic fields in the etiology of childhood leukemia becomes more evident.

Adolescent↗

Risk factors of delayed extubation, prolonged length of stay in the intensive care unit, and mortality in patients undergoing coronary artery bypass graft with fast-track cardiac anesthesia: a new cardiac risk score.

BACKGROUND: Risk factors of delayed extubation, prolonged intensive care unit (ICU) length of stay (LOS), and mortality have not been studied for patients administered fast-track cardiac anesthesia (FTCA). The authors' goals were to determine risk factors of outcomes and cardiac risk scores (CRS) for CABG patients undergoing FTCA. METHODS: Consecutive CABG patients undergoing FTCA were prospectively studied. Outcome variables were delayed extubation > 10 h, prolonged ICU LOS > 48 h, and mortality. Univariate analyses were performed followed by multiple logistic regression to derive risk factors of the three outcomes. Simplified integer-based CRS were derived from logistic models. Bootstrap validation was performed to assess and compare the predictive abilities of CRS and logistic models for the three outcomes. RESULTS: The authors studied 885 patients. Twenty-five percent had delayed extubation, 17% had prolonged ICU LOS, and 2.6% died. Risk factors of delayed extubation were increased age, female gender, postoperative use of intraaortic balloon pump, inotropes, bleeding, and atrial arrhythmia. Risk factors of prolonged ICU LOS were those of delayed extubation plus preoperative myocardial infarction and postoperative renal insufficiency. Risk factors of mortality were female gender, emergency surgery, and poor left ventricular function. CRSs were modeled for the three outcomes. The area under the receiver operating characteristic curve for the CRS-logistic models was not significantly different: 0.707/0.702 for delayed extubation, 0.851/0.855 for prolonged ICU LOS, and 0.657/0.699 for mortality. CONCLUSION: In CABG patients undergoing FTCA, the authors derived and validated risk factors of delayed extubation, prolonged ICU LOS, and mortality. Furthermore, they developed a simplified CRS system with similar predictive abilities as the logistic models.

Aged↗

A comparison of statistical learning methods on the Gusto database.

We apply a battery of modern, adaptive non-linear learning methods to a large real database of cardiac patient data. We use each method to predict 30 day mortality from a large number of potential risk factors, and we compare their performances. We find that none of the methods could outperform a relatively simple logistic regression model previously developed for this problem.

Databases as Topic↗

Impact of menstrual phase on false-negative mammograms in the Canadian National Breast Screening Study.

BACKGROUND: The efficacy of breast carcinoma screening should be enhanced if false-negative mammography were reduced. Prospectively collected data from the Canadian National Breast Screening Study were used to examine whether menstrual cycle phase was associated with false-negative outcomes for mammographic screening. METHODS: Of 8887 women ages 40-44 years at the onset of screening, randomized to receive annual mammography and clinical breast examination, reporting menstruation no more than 28 days prior to their screening examination, and with a valid radiologic report, 1898 had never used oral contraceptives or replacement estrogen with or without progesterone. The remainder were past (6573) and current (416) estrogen users. Similar selection criteria were applied at subsequent screens. The distribution of false-negative and false-positive mammography in relation to true-negative and true-positive mammography was examined with respect to the follicular (Days 1 to 14) and luteal (Days 15-28) menstrual phases. RESULTS: Comparing luteal with follicular mammograms in 6989 patients who ever used estrogen, the unadjusted odds ratio (2-sided P-values) for false-negatives versus true-negatives was 2.16 (0.05) and the adjusted odds ratio was 1.47 (0.05). In 1898 never-users, parallel odds ratios for luteal false-negatives were 0.55 (1.0) and 0.74 (1.0), respectively. CONCLUSIONS: These results suggest that menstruating women who have used hormones may have an increased risk of false-negative results for screening mammograms performed in the luteal phase of the menstrual cycle. An increased risk of false-negative mammography might adversely affect screening efficacy. The impact of menstrual phase on mammographic interpretation, especially for women who ever used hormones, requires further investigation.

Adult↗

The lasso method for variable selection in the Cox model.

I propose a new method for variable selection and shrinkage in Cox's proportional hazards model. My proposal minimizes the log partial likelihood subject to the sum of the absolute values of the parameters being bounded by a constant. Because of the nature of this constraint, it shrinks coefficients and produces some coefficients that are exactly zero. As a result it reduces the estimation variance while providing an interpretable final model. The method is a variation of the 'lasso' proposal of Tibshirani, designed for the linear regression context. Simulations indicate that the lasso can be more accurate than stepwise selection in this setting.

Humans↗

Generalized additive models for medical research.

This article reviews flexible statistical methods that are useful for characterizing the effect of potential prognostic factors on disease endpoints. Applications to survival models and binary outcome models are illustrated.

Algorithms↗

Detection of silent ischemia adds to the prognostic value of coronary anatomy and left ventricular function in predicting outcome in unstable angina patients.

BACKGROUND: Patients with unstable angina are at increased risk of unfavourable outcomes such as myocardial infarction, death and urgent revascularization. Early risk stratification may improve subsequent outcome. Recently the presence and duration (at least 60 mins) of silent ischemia as measured by Holter monitoring has been shown to be of prognostic value. The incremental value of this information over that provided by coronary angiography and assessment of left ventricular function is not known. OBJECTIVE: To determine whether detection of silent ischemia is of independent and additional prognostic significance beyond that provided by the angiographic extent of coronary artery disease and left ventricular dysfunction. METHODS: One hundred and thirty-five unstable angina patients with 24 h of ST segment monitoring in addition to early cardiac catheterization (4 +/- 3 days) were assessed. Eighty-nine patients (66%) had ST segment shift for a total of 593 episodes (mean duration of 18 +/- 30 mins per episode) of which 92% were asymptomatic. Ten patients had a myocardial infarction and six patients died during the hospitalization. In addition, there were 33 urgent revascularization procedures. RESULTS: With the generalized additive logistic model, various clinical variables were assessed for predicting unfavourable outcomes. Duration of ST shift (P = 0.02) was second only to angiographic severity of coronary artery disease (P = 0.004) as a predictor. In the presence of these two variables left ventricular function did not have independent prognostic significance (P = 0.16). Event-free survival curves show that duration of ST shift of at least 60 mins was of incremental value in predicting unfavourable in-hospital outcomes compared with both the extent of coronary artery disease and left ventricular dysfunction. CONCLUSION: In patients with unstable angina, further stratification can be achieved early with Holter monitoring in addition to coronary angiography and assessment of left ventricular function.

Angina, Unstable↗

Implications of measurement error in exposure for the sample sizes of case-control studies.

In this paper, recent results describing the effects of measurement error on estimation of the association between an exposure and a disease are applied to sample size calculation in case-control studies. Models of the relation between true exposure and a surrogate exposure measure assessed with error are used to derive equations for sample size determination. The results show that the sample size of a study based on an exposure variable which is measured with error must be larger by a factor of 1/rho 2 than if exposure were measured without error, where rho is the correlation between the true exposure and the surrogate exposure measure. Review of the magnitude of measurement error in dietary assessments suggests that failure to account for measurement error in sample size determination for case-control studies of diet and disease could lead to marked underestimation of the required sample size.

Bias↗

Spinal deformity after multiple-level cervical laminectomy in children.

Considerable controversy exists in the orthopedic and neurosurgical literature over the true incidence and nature of spinal deformity after multiple-level cervical laminectomy in children. Eighty-nine patients with a mean radiographic follow up of 5.1 years (range 2-9 years) were reviewed. Mean age at surgery was 5.7 years (range 1 month-18 years). Most common diagnoses were Arnold-Chiari malformation, syringomyelia, or both (81%). Significant deformity developed in 46 patients (53%), with 33 developing a mean kyphosis of 30 degrees (range 5-105 degrees) and 13 developing a mean hyperlordosis of 62 degrees (range 40-95 degrees). Peak age at surgery of 10.5 years correlated weakly (P = 0.08) with the development of kyphosis. The development of hyperlordosis was strongly correlated (P = 0.01) with a peak age at surgery of 4.2 years. There was no correlation between diagnosis, sex, location, or number of levels decompressed and the subsequent development of deformity.

Adolescent↗