Search PubMed⌕ Search

Biomedical subjects

Hongyue Dai

Publications and source records attributed to Hongyue Dai.

17 recordsLinked to original sources

Rosetta error model for gene expression analysis.

MOTIVATION: In microarray gene expression studies, the number of replicated microarrays is usually small because of cost and sample availability, resulting in unreliable variance estimation and thus unreliable statistical hypothesis tests. The unreliable variance estimation is further complicated by the fact that the technology-specific variance is intrinsically intensity-dependent. RESULTS: The Rosetta error model captures the variance-intensity relationship for various types of microarray technologies, such as single-color arrays and two-color arrays. This error model conservatively estimates intensity error and uses this value to stabilize the variance estimation. We present two commonly used error models: the intensity error-model for single-color microarrays and the ratio error model for two-color microarrays or ratios built from two single-color arrays. We present examples to demonstrate the strength of our error models in improving statistical power of microarray data analysis, particularly, in increasing expression detection sensitivity and specificity when the number of replicates is limited.

Algorithms↗

Gene expression changes associated with progression and response in chronic myeloid leukemia.

Chronic myeloid leukemia (CML) is a hematopoietic stem cell disease with distinct biological and clinical features. The biologic basis of the stereotypical progression from chronic phase through accelerated phase to blast crisis is poorly understood. We used DNA microarrays to compare gene expression in 91 cases of CML in chronic (42 cases), accelerated (17 cases), and blast phases (32 cases). Three thousand genes were found to be significantly (P < 10(-10)) associated with phase of disease. A comparison of the gene signatures of chronic, accelerated, and blast phases suggest that the progression of chronic phase CML to advanced phase (accelerated and blast crisis) CML is a two-step rather than a three-step process, with new gene expression changes occurring early in accelerated phase before the accumulation of increased numbers of leukemia blast cells. Especially noteworthy and potentially significant in the progression program were the deregulation of the WNT/beta-catenin pathway, the decreased expression of Jun B and Fos, alternative kinase deregulation, such as Arg (Abl2), and an increased expression of PRAME. Studies of CML patients who relapsed after initially successful treatment with imatinib demonstrated a gene expression pattern closely related to advanced phase disease. These studies point to specific gene pathways that might be exploited for both prognostic indicators as well as new targets for therapy.

Antineoplastic Agents↗

A cell proliferation signature is a marker of extremely poor outcome in a subpopulation of breast cancer patients.

Breast cancer comprises a group of distinct subtypes that despite having similar histologic appearances, have very different metastatic potentials. Being able to identify the biological driving force, even for a subset of patients, is crucially important given the large population of women diagnosed with breast cancer. Here, we show that within a subset of patients characterized by relatively high estrogen receptor expression for their age, the occurrence of metastases is strongly predicted by a homogeneous gene expression pattern almost entirely consisting of cell cycle genes (5-year odds ratio of metastasis, 24.0; 95% confidence interval, 6.0-95.5). Overexpression of this set of genes is clearly associated with an extremely poor outcome, with the 10-year metastasis-free probability being only 24% for the poor group, compared with 85% for the good group. In contrast, this gene expression pattern is much less correlated with the outcome in other patient subpopulations. The methods described here also illustrate the value of combining clinical variables, biological insight, and machine-learning to dissect biological complexity. Our work presented here may contribute a crucial step towards rational design of personalized treatment.

Age Factors↗

A protocol for building and evaluating predictors of disease state based on microarray data.

MOTIVATION: Microarray gene expression data are increasingly employed to identify sets of marker genes that accurately predict disease development and outcome in cancer. Many computational approaches have been proposed to construct such predictors. However, there is, as yet, no objective way to evaluate whether a new approach truly improves on the current state of the art. In addition no 'standard' computational approach has emerged which enables robust outcome prediction. RESULTS: An important contribution of this work is the description of a principled training and validation protocol, which allows objective evaluation of the complete methodology for constructing a predictor. We review the possible choices of computational approaches, with specific emphasis on predictor choice and reporter selection strategies. Employing this training-validation protocol, we evaluated different reporter selection strategies and predictors on six gene expression datasets of varying degrees of difficulty. We demonstrate that simple reporter selection strategies (forward filtering and shrunken centroids) work surprisingly well and outperform partial least squares in four of the six datasets. Similarly, simple predictors, such as the nearest mean classifier, outperform more complex classifiers. Our training-validation protocol provides a robust methodology to evaluate the performance of new computational approaches and to objectively compare outcome predictions on different datasets.

Algorithms↗

Robustness, scalability, and integration of a wound-response gene expression signature in predicting breast cancer survival.

Based on the hypothesis that features of the molecular program of normal wound healing might play an important role in cancer metastasis, we previously identified consistent features in the transcriptional response of normal fibroblasts to serum, and used this "wound-response signature" to reveal links between wound healing and cancer progression in a variety of common epithelial tumors. Here, in a consecutive series of 295 early breast cancer patients, we show that both overall survival and distant metastasis-free survival are markedly diminished in patients whose tumors expressed this wound-response signature compared to tumors that did not express this signature. A gene expression centroid of the wound-response signature provides a basis for prospectively assigning a prognostic score that can be scaled to suit different clinical purposes. The wound-response signature improves risk stratification independently of known clinico-pathologic risk factors and previously established prognostic signatures based on unsupervised hierarchical clustering ("molecular subtypes") or supervised predictors of metastasis ("70-gene prognosis signature").

Breast Neoplasms↗

Diet induction of monocyte chemoattractant protein-1 and its impact on obesity.

OBJECTIVE: To examine the effect of a high-fat diet on gene expression in adipose tissues and to determine induction kinetics of adipose monocyte chemoattractant protein-1 and -3 (MCP-1 and MCP-3) in diet-induced obesity (DIO) and the effect of a lack of MCP-1 signaling on DIO susceptibility and macrophage recruitment into adipose tissue. RESEARCH METHODS AND PROCEDURES: Obese and lean adipose tissues were profiled for expression changes. The time-course of MCP-1 and MCP-3 expression was examined by reverse transcriptase-polymerase chain reaction. Plasma MCP-1 levels were determined by enzyme-linked immunosorbent assay (ELISA). Chemokine receptor-2 (CCR2) knockout mice were placed on the high-fat diet to determine DIO susceptibility. Macrophage infiltration in adipose tissue was examined by immunohistochemistry with F4/80 antibody. RESULTS: DIO elevated adipose expression of many inflammatory genes, including MCP-1 and MCP-3. Adipose MCP-1 and MCP-3 mRNA levels increased within 7 days of starting a high-fat diet, with elevation of plasma MCP-1 detected after 4 weeks on the diet. The induction of MCP-1 and MCP-3 expression preceded that of tumor necrosis factor-alpha. The elevated plasma MCP-1 concentration in obese mice was partially reversed by treatment with AM251. No change in DIO susceptibility and macrophage accumulation in adipose tissue were observed in CCR2 knockout mice, which lack the MCP-1 receptor CCR2. DISCUSSION: A high-fat diet elevated adipose expression of inflammatory genes, including early induction of MCP-1 and MCP-3, supporting the view that obese adipose tissues contribute to systemic inflammation. However, despite increased MCP-1 in obesity, disruption of MCP-1 signaling did not confer resistance to DIO in mice or reduce adipose tissue macrophage infiltration.

Adipocytes↗

Identification of biomarkers for tumor endothelial cell proliferation through gene expression profiling.

Extensive efforts are under way to identify antiangiogenic therapies for the treatment of human cancers. Many proposed therapeutics target vascular endothelial growth factor (VEGF) or the kinase insert domain receptor (KDR/VEGF receptor-2/FLK-1), the mitogenic VEGF receptor tyrosine kinase expressed by endothelial cells. Inhibition of KDR catalytic activity blocks tumor neoangiogenesis, reduces vascular permeability, and, in animal models, inhibits tumor growth and metastasis. Using a gene expression profiling strategy in rat tumor models, we identified a set of six genes that are selectively overexpressed in tumor endothelial cells relative to tumor cells and whose pattern of expression correlates with the rate of tumor endothelial cell proliferation. In addition to being potential targets for antiangiogenesis tumor therapy, the expression patterns of these genes or their protein products may aid the development of pharmacodynamic assays for small molecule inhibitors of the KDR kinase in human tumors.

Angiogenesis Inhibitors↗

Integrated genomic and proteomic analyses of gene expression in Mammalian cells.

Using DNA microarrays together with quantitative proteomic techniques (ICAT reagents, two-dimensional DIGE, and MS), we evaluated the correlation of mRNA and protein levels in two hematopoietic cell lines representing distinct stages of myeloid differentiation, as well as in the livers of mice treated for different periods of time with three different peroxisome proliferative activated receptor agonists. We observe that the differential expression of mRNA (up or down) can capture at most 40% of the variation of protein expression. Although the overall pattern of protein expression is similar to that of mRNA expression, the incongruent expression between mRNAs and proteins emphasize the importance of posttranscriptional regulatory mechanisms in cellular development or perturbation that can be unveiled only through integrated analyses of both proteins and mRNAs.

Animals↗

A microarray platform comparison for neuroscience applications.

To address the need for high sensitivity in gene expression profiling of small neural tissue samples ( approximately 100 ng total RNA), we compared a novel RT-PCR-IVT protocol using fluor-reverse pairs on inkjet oligonucleotide microarrays and an RT-IVT protocol using 33P labeling on nylon cDNA arrays. The comparison protocol was designed to evaluate these systems for sensitivity, specificity, reproducibility, and linearity. We developed parameters, thresholds, and testing conditions that could be used to differentiate various systems that spanned detection chemistry and instrumentation; probe number and selection criteria; and sample processing protocols. We concluded that the inkjet system had better performance in sensitivity, specificity, and reproducibility than the nylon system, and similar performance in linearity. Between these two platforms, the data indicates that the inkjet system would perform better for the transcriptional profiling of 100 ng total RNA samples for neuroscience studies.

Animals↗

T lymphocyte activation gene identification by coregulated expression on DNA microarrays.

High-capacity methods for assessing gene function have become increasingly important because of the increasing number of newly identified genes emerging from large-scale genome sequencing and cDNA cloning efforts. We investigated the use of DNA microarrays to identify uncharacterized genes specifically involved in human T cell activation. Activation of human peripheral blood T lymphocytes induced significant changes in hundreds of transcripts, but most of these were not unique to T cell activation. Variation of experimental parameters and analysis techniques allowed better enrichment for gene expression changes unique to T cell activation. Best results were achieved by identification of genes that were most highly coregulated with the T-cell-specific transcript interleukin 2 (IL2) in a "compendium" of experiments involving both T cells and other cell types. Among the genes most highly coregulated with IL2 were many genes known to function during T cell activation, together with ESTs of unknown function. Four of these ESTs were extended to novel full-length clones encoding T-cell-regulated proteins with predicted functions in GTP metabolism, cell organization, and signal transduction.

Base Sequence↗

Signatures of environmental exposures using peripheral leukocyte gene expression: tobacco smoke.

Functional biological markers of environmental exposures are important in epidemiological studies of disease risk. Such markers not only provide a measure of the exposure, they also reflect the degree of physiological and biochemical response to the exposure. In an observational study, using DNA microarrays, we show that it is possible to distinguish between 85 individuals exposed and unexposed to tobacco smoke on the basis of mRNA expression in peripheral leukocytes. Furthermore, we show that active exposure to tobacco smoke is associated with a biologically relevant mRNA expression signature. These findings suggest that expression patterns can be used to identify a complex environmental exposure in humans.

Adult↗

Effects of atmospheric ozone on microarray data quality.

A data anomaly was observed that affected the uniformity and reproducibility of fluorescent signal across DNA microarrays. Results from experimental sets designed to identify potential causes (from microarray production to array scanning) indicated that the anomaly was linked to a batch process; further work allowed us to localize the effect to the posthybridization array stringency washes. Ozone levels were monitored and highly correlated with the batch effect. Controlled exposures of microarrays to ozone confirmed this factor as the root cause, and we present data that show susceptibility of a class of cyanine dyes (e.g., Cy5, Alexa 647) to ozone levels as low as 5-10 ppb for periods as short as 10-30 s. Other cyanine dyes (e.g., Cy3, Alexa 555) were not significantly affected until higher ozone levels (> 100 ppb). To address this environmental effect, laboratory ozone levels should be kept below 2 ppb (e.g., with filters in HVAC) to achieve high quality microarray data.

Artifacts↗

Microarray standard data set and figures of merit for comparing data processing methods and experiment designs.

MOTIVATION: There is a very large and growing level of effort toward improving the platforms, experiment designs, and data analysis methods for microarray expression profiling. Along with a growing richness in the approaches there is a growing confusion among most scientists as to how to make objective comparisons and choices between them for different applications. There is a need for a standard framework for the microarray community to compare and improve analytical and statistical methods. RESULTS: We report on a microarray data set comprising 204 in-situ synthesized oligonucleotide arrays, each hybridized with two-color cDNA samples derived from 20 different human tissues and cell lines. Design of the approximately 24 000 60mer oligonucleotides that report approximately 2500 known genes on the arrays, and design of the hybridization experiments, were carried out in a way that supports the performance assessment of alternative data processing approaches and of alternative experiment and array designs. We also propose standard figures of merit for success in detecting individual differential expression changes or expression levels, and for detecting similarities and differences in expression patterns across genes and experiments. We expect this data set and the proposed figures of merit will provide a standard framework for much of the microarray community to compare and improve many analytical and statistical methods relevant to microarray data analysis, including image processing, normalization, error modeling, combining of multiple reporters per gene, use of replicate experiments, and sample referencing schemes in measurements based on expression change. AVAILABILITY/SUPPLEMENTARY INFORMATION: Expression data and supplementary information are available at http://www.rii.com/publications/2003/HE_SDS.htm

Base Sequence↗

A gene-expression signature as a predictor of survival in breast cancer.

BACKGROUND: A more accurate means of prognostication in breast cancer will improve the selection of patients for adjuvant systemic therapy. METHODS: Using microarray analysis to evaluate our previously established 70-gene prognosis profile, we classified a series of 295 consecutive patients with primary breast carcinomas as having a gene-expression signature associated with either a poor prognosis or a good prognosis. All patients had stage I or II breast cancer and were younger than 53 years old; 151 had lymph-node-negative disease, and 144 had lymph-node-positive disease. We evaluated the predictive power of the prognosis profile using univariable and multivariable statistical analyses. RESULTS: Among the 295 patients, 180 had a poor-prognosis signature and 115 had a good-prognosis signature, and the mean (+/-SE) overall 10-year survival rates were 54.6+/-4.4 percent and 94.5+/-2.6 percent, respectively. At 10 years, the probability of remaining free of distant metastases was 50.6+/-4.5 percent in the group with a poor-prognosis signature and 85.2+/-4.3 percent in the group with a good-prognosis signature. The estimated hazard ratio for distant metastases in the group with a poor-prognosis signature, as compared with the group with the good-prognosis signature, was 5.1 (95 percent confidence interval, 2.9 to 9.0; P<0.001). This ratio remained significant when the groups were analyzed according to lymph-node status. Multivariable Cox regression analysis showed that the prognosis profile was a strong independent factor in predicting disease outcome. CONCLUSIONS: The gene-expression profile we studied is a more powerful predictor of the outcome of disease in young patients with breast cancer than standard systems based on clinical and histologic criteria.

Adult↗

Use of hybridization kinetics for differentiating specific from non-specific binding to oligonucleotide microarrays.

Hybridization kinetics were found to be significantly different for specific and non-specific binding of labeled cRNA to surface-bound oligonucleotides on microarrays. We show direct evidence that in a complex sample specific binding takes longer to reach hybridization equilibrium than the non- specific binding. We find that this property can be used to estimate and to correct for the hybridization contributed by non-specific binding. Useful applications are illustrated including the selection of superior oligonucleotides, and the reduction of false positives in exon identification.

Base Pair Mismatch↗

Gene expression profiling predicts clinical outcome of breast cancer.

Breast cancer patients with the same stage of disease can have markedly different treatment responses and overall outcome. The strongest predictors for metastases (for example, lymph node status and histological grade) fail to classify accurately breast tumours according to their clinical behaviour. Chemotherapy or hormonal therapy reduces the risk of distant metastases by approximately one-third; however, 70-80% of patients receiving this treatment would have survived without it. None of the signatures of breast cancer gene expression reported to date allow for patient-tailored therapy strategies. Here we used DNA microarray analysis on primary breast tumours of 117 young patients, and applied supervised classification to identify a gene expression signature strongly predictive of a short interval to distant metastases ('poor prognosis' signature) in patients without tumour cells in local lymph nodes at diagnosis (lymph node negative). In addition, we established a signature that identifies tumours of BRCA1 carriers. The poor prognosis signature consists of genes regulating cell cycle, invasion, metastasis and angiogenesis. This gene expression profile will outperform all currently used clinical parameters in predicting disease outcome. Our findings provide a strategy to select patients who would benefit from adjuvant therapy.

Adult↗