Search PubMed⌕ Search

Biomedical subjects

Shuguang Huang

Publications and source records attributed to Shuguang Huang.

13 recordsLinked to original sources

Comparative analysis and integrative classification of NCI60 cell lines and primary tumors using gene expression profiling data.

BACKGROUND: NCI60 cell lines are derived from cancers of 9 tissue origins and have been invaluable in vitro models for cancer research and anti-cancer drug screen. Although extensive studies have been carried out to assess the molecular features of NCI60 cell lines related to cancer and their sensitivities to more than 100,000 chemical compounds, it remains unclear if and how well these cell lines represent or model their tumor tissues of origin. Identification and confirmation of correct origins of NCI60 cell lines are critical to their usage as model systems and to translate in vitro studies into clinical potentials. Here we report a direct comparison between NCI60 cell lines and primary tumors by analyzing global gene expression profiles. RESULTS: Comparative analysis suggested that 51 of 59 cell lines we analyzed represent their presumed tumors of origin. Taking advantage of available clinical information of primary tumor samples used to generate gene expression profiling data, we further classified those cell lines with the correct origins into different subtypes of cancer or different stages in cancer development. For example, 6 of 7 non-small cell lung cancer cell lines were classified as lung adenocarcinomas and all of them were classified into late stages in tumor progression. CONCLUSION: Taken together, we developed and applied a novel approach for systematic comparative analysis and integrative classification of NCI60 cell lines and primary tumors. Our results could provide guidance to the selection of appropriate cell lines for cancer research and pharmaceutical compound screenings. Moreover, this gene expression profile based approach can be generally applied to evaluate experimental model systems such as cell lines and animal models for human diseases.

Carcinoma, Non-Small-Cell Lung↗

The loss in power when the test of differential expression is performed under a wrong scale.

One of the most common and important goals of microarray studies is to identify genes that are differentially expressed between cells of different conditions. T-test and ANOVA models on the expression data are common practices to gauge the significance of the observed difference in expression levels. Transformation of the microarray data is often applied in order to satisfy the model assumptions being entertained. However, the distributional properties of the expression are gene specific, and it is impractical to find a single transformation that is universally optimal for all the genes. This difficulty results in the situation that some genes have to violate the assumptions of the model (e.g., homogeneity in variance, normality). It is thus the interest of this paper to evaluate the impact on the inference of differential expression when the test is performed under an inappropriate scale. Particularly, we quantitatively assess the loss of power when the test is performed under a wrong scale. Normal distribution and log-normal distribution of the expression data are considered. The loss in power is investigated in two scenarios: a transformation is misused, or a transformation fails to be applied. Log transformation and power transformation are particularly considered due to the fact that Box-Cox types of transformation are commonly used in practice. The impact of using a wrong scale is investigated analytically and based on simulations. The loss in power is assessed both as a function of the degree to which the assumptions are violated and as a function of the effect size. Simulations are conducted to quantitatively assess the power loss when tests are performed under a wrong scale. A public experimental microarray dataset is used to illustrate the impact of transformation on the results of testing differential expression. The results show that the loss of power is a function of CV and fold-change (effect size). The loss in power depends on the true model and on how severely the assumptions are violated.

Databases, Genetic↗

Expression profiling of rat femur revealed suppression of bone formation genes by treatment with alendronate and estrogen but not raloxifene.

The pharmacological preservation of bone in the ovariectomized rat by estrogen, selective estrogen receptor modulators (SERMs), and bisphosphonates has been well described. However, comprehensive molecular analysis of the effects of these pharmacologically diverse antiresorptive agents on gene expression in bone has not been performed. This study used DNA microarrays to analyze RNA from the proximal femur metaphysis of sham and ovariectomized vehicle-treated rats, and ovariectomized rats treated for 35 days with maximally efficacious doses of 17-alpha ethinyl estradiol, the benzothiophene SERM, raloxifene, the benzopyran SERM, (S)-3-(4-hydroxyphenyl)-4-methyl-2-[4-[2-(1-piperidinyl)ethoxy]phenyl]-2H-1-benzopyran-7-ol (EM652), and the aminobisphosphonate, alendronate. Ovariectomy resulted in 644 significant probe set changes relative to sham control rats (p < 0.05), whereas E2, raloxifene, EM652, and alendronate regulated 613, 765, 652, and 737 probe sets, respectively, relative to ovariectomized control rats. An intersection of these data sets yielded 334 unique genes that were altered after ovariectomy and additionally changed by one or more antiresorptive treatment. Clustering analysis showed that the transcript profile was distinctly different for each pharmaceutical agent and that raloxifene maintained more genes at sham levels than any other treatment. In addition, E2 and alendronate suppressed a cluster of genes associated with bone formation activity below that of sham, whereas raloxifene had little effect on these genes. These data indicate stronger suppressive effects of E2 and alendronate on bone formation activity and that ovariectomy plus raloxifene resembles sham more closely than ovariectomized animals treated with E2, EM652, or alendronate.

Alendronate↗

Pooling samples within microarray studies: a comparative analysis of rat liver transcription response to prototypical toxicants.

Combining or pooling individual samples when carrying out transcript profiling using microarrays is a fairly common means to reduce both the cost and complexity of data analysis. However, pooling does not allow for statistical comparison of changes between samples and can result in a loss of information. Because a rigorous comparison of the identified expression changes from the two approaches has not been reported, we compared the results for hepatic transcript profiles from pooled vs. individual samples. Hepatic transcript profiles from a single-dose time-course rat study in response to the prototypical toxicants clofibrate, diethylhexylphthalate, and valproic acid were evaluated. Approximately 50% more transcript expression changes were observed in the individual (statistical) analysis compared with the pooled analysis. While the majority of these changes were less than twofold in magnitude ( approximately 80%), a substantial number were greater than twofold (approximately 20%). Transcript changes unique to the individual analysis were confirmed by quantitative RT-PCR, while all the changes unique to the pooled analysis did not confirm. The individual analysis identified more hits per biological pathway than the pooled approach. Many of the transcripts identified by the individual analysis were novel findings and may contribute to a better understanding of molecular mechanisms of these compounds. Furthermore, having individual animal data provided the opportunity to correlate changes in transcript expression to phenotypes (i.e., histology) observed in toxicology studies. The two approaches were similar when clustering methods were used despite the large difference in the absolute number of transcripts changed. In summary, pooling reduced resource requirements substantially, but the individual approach enabled statistical analysis that identified more gene expression changes to evaluate mechanisms of toxicity. An individual animal approach becomes more valuable when the overall expression response is subtle and/or when associating expression data to variable phenotypic responses.

Animals↗

Molecular profile of catabolic versus anabolic treatment regimens of parathyroid hormone (PTH) in rat bone: an analysis by DNA microarray.

Teriparatide, human PTH (1-34), a new therapy for osteoporosis, elicits markedly different skeletal responses depending on the treatment regimen. In order to understand potential mechanisms for this dichotomy, the present investigation utilized microarrays to delineate the genes and pathways that are regulated by intermittent (subcutaneous injection of 80 microg/kg/day) and continuous (subcutaneous infusion of 40 microg/kg/day by osmotic mini pump) PTH (1-34) for 1 week in 6-month-old female rats. The effect of each PTH regimen was confirmed by histomorphometric analysis of the proximal tibial metaphysis, and mRNA from the distal femoral metaphysis was analyzed using an Affymetrix microarray. Both PTH paradigms co-regulated 22 genes including known bone formation genes (i.e., collagens, osteocalcin, decorin, and osteonectin) and also uniquely modulated additional genes. Intermittent PTH regulated 19 additional genes while continuous treatment regulated 173 additional genes. This investigation details for the first time the broad profiling of the gene and pathway changes that occur in vivo following treatment of intermittent versus continuous PTH (1-34). These results extend previous observations of gene expression changes and reveal the in vivo regulation of BMP3 and multiple neuronal genes by PTH treatment.

Animals↗

Comparison of false discovery rate methods in identifying genes with differential expression.

Current high-throughput techniques such as microarray in genomics or mass spectrometry in proteomics usually generate thousands of hypotheses to be tested simultaneously. The usual purpose of these techniques is to identify a subset of interesting cases that deserve further investigation. As a consequence, the control of false positives among the tests called "significant" becomes a critical issue for researchers. Over the past few years, several false discovery rate (FDR)-controlling methods have been proposed; each method favors certain scenarios and is introduced with the purpose of improving the control of FDR at the targeted level. In this paper, we compare the performance of the five FDR-controlling methods proposed by Benjamini et al., the qvalue method proposed by Storey, and the traditional Bonferroni method. The purpose is to investigate the "observed" sensitivity of each method on typical microarray experiments in which the majority (or all) of the truth is unknown. Based on two well-studied microarray datasets, it is found that in terms of the "apparent" test power, the ranking of the FDR methods is given as Step-down<Step-up: dependent<Step-up: one-stage (BH95)<Step-up adaptive<qvalue. The BH95 method shows the best control of FDR at the target level. It is our hope that the observed results could provide some insight into the application of different FDR methods in microarray data analysis.

Algorithms↗

Transcriptional profiling of keratinocytes reveals a vitamin D-regulated epidermal differentiation network.

1alpha,25-dihydroxyvitamin D(3) [1,25(OH)(2)D(3)] regulates mineral homeostasis and exhibits potent anti-proliferative, prodifferentiative, and immunomodulatory activities. It mediates these effects by binding to the vitamin D receptor (VDR), which belongs to the superfamily of steroid/thyroid hormone nuclear receptors. As a result of keratinocyte differentiation and anti-proliferation activities, 1,25(OH)(2)D(3) and its synthetic analogs are therapeutically effective in psoriasis and show promise for the treatment of actinic keratosis and squamous cell carcinoma. To elucidate the VDR signaling pathway in keratinocytes, we examined the gene expression profile with 1,25(OH)(2)D(3) treatment using oligonucleotide microarrays. Out of the 12,600 genes investigated, 82 were upregulated and 16 were downregulated and many of these were involved in differentiation, proliferation, and immune response. We have identified three vitamin D-responsive chromosomal loci (1p36, 19q13, and 6p25) and show the induction of various class II tumor suppressor/growth-regulatory genes in response to 1,25(OH)(2)D(3). Finally, quantitative differences in gene expression revealed a vitamin D-regulated differentiation network and identified peptidylarginine deiminases, kallikreins, serine proteinase inhibitor family members, Kruppel-like factor 4, and c-fos as vitamin D-responsive genes, whose protein products may play an important role in epidermal differentiation in normal and diseased state.

Calcitriol↗

Overexpression of G protein-coupled receptors in cancer cells: involvement in tumor progression.

G protein-coupled receptors (GPCRs) play important roles in a variety of biological and pathological processes. They are considered among the most desirable targets for drug development. Recent studies have demonstrated that many GPCRs, such as endothelin receptors, chemokine receptors and lysophosphatidic acid receptors have been implicated in the tumorigenesis and metastasis of multiple human cancers. In this study, we conducted an in silico analysis of GPCR gene expression in primary human tumors by analyzing some publicly available gene expression profiling data. Statistical analysis was performed on eight microarray data sets of non-small cell lung cancer, breast cancer, prostate cancer, melanoma, gastric cancer and diffused large B cell lymphoma to identify GPCRs that are up-regulated in primary or metastatic cancer cells. Our analysis has demonstrated overexpression of several GPCRs in primary tumor cells, including chemokine receptors and protease-activated receptors that were shown to be important for tumorigenesis by previous studies. In addition, we have uncovered several GPCRs, such as neuropeptide receptors, adenosine A2B receptor, P2Y purinoceptor, calcium-sensing receptor and metabotropic glutamate receptors, that are expressed at a significantly higher level in some cancer tissue and may play a role in cancer progression. Analysis of cancer samples in different disease stages also suggests that some GPCRs, such as endothelin receptor A, may be involved in early tumor progression and others, such as CXCR4, may play a critical role in tumor invasion and metastasis. The present study demonstrates the value of publicly available microarray data as a resource to gain more understanding of cancer biology, to validate previous findings from in vitro experiments, and to identify potential novel anticancer targets and biomarkers.

Biomarkers, Tumor↗

SUM: a new way to incorporate mismatch probe measurements.

Affymetrix's high-density oligonucleotide arrays offer an exciting technology in biomedical research. With more and more statistical involvement in every step of the process, there has been a constant effort to make sure that the expression data are appropriately extracted in the first place. According to Affymetrix GeneChip technology, each gene is represented by 11-20 oligo probe pairs; the challenge is how to extract one meaningful number, expression, from the 11-20 pairs of numbers. More specifically, there is first a need to differentiate the components of specific binding, nonspecific binding, and optical background noise in both PM and MM probes, and then an expression measure that is proportional to the true abundance of transcripts is to be derived. A new method, SUM, which sums up PM and MM values and then follows a process similar to that of RMA, is considered. The performance of SUM is investigated and compared to the three most popular methods, MAS5, dChip, and RMA. The assessments are based on a well-controlled experiment dataset that is publicly available. The results show that in several respects the performance of SUM is comparable to that of RMA and dChip, and all three of these methods show some advantages over MAS5. There is some evidence showing that SUM has higher differential sensitivity than other methods in certain situations.

Algorithms↗

At what scale should microarray data be analyzed?

INTRODUCTION: The hybridization intensities derived from microarray experiments, for example Affymetrix's MAS5 signals, are very often transformed in one way or another before statistical models are fitted. The motivation for performing transformation is usually to satisfy the model assumptions such as normality and homogeneity in variance. Generally speaking, two types of strategies are often applied to microarray data depending on the analysis need: correlation analysis where all the gene intensities on the array are considered simultaneously, and gene-by-gene ANOVA where each gene is analyzed individually. AIM: We investigate the distributional properties of the Affymetrix GeneChip signal data under the two scenarios, focusing on the impact of analyzing the data at an inappropriate scale. METHODS: The Box-Cox type of transformation is first investigated for the strategy of pooling genes. The commonly used log-transformation is particularly applied for comparison purposes. For the scenario where analysis is on a gene-by-gene basis, the model assumptions such as normality are explored. The impact of using a wrong scale is illustrated by log-transformation and quartic-root transformation. RESULTS: When all the genes on the array are considered together, the dependent relationship between the expression and its variation level can be satisfactorily removed by Box-Cox transformation. When genes are analyzed individually, the distributional properties of the intensities are shown to be gene dependent. Derivation and simulation show that some loss of power is incurred when a wrong scale is used, but due to the robustness of the t-test, the loss is acceptable when the fold-change is not very large.

Algorithms↗

Identification of genes for complex disease using longitudinal phenotypes.

Using the simulated data set from Genetic Analysis Workshop 13, we explored the advantages of using longitudinal data in genetic analyses. The weighted average of the longitudinal data for each of seven quantitative phenotypes were computed and analyzed. Genome screen results were then compared for these longitudinal phenotypes and the results obtained using two cross-sectional designs: data collected near a single age (45 years) and data collected at a single time point. Significant linkage was obtained for nine regions (LOD scores ranging from 5.5 to 34.6) for six of the phenotypes. Using cross-sectional data, LOD scores were slightly lower for the same chromosomal regions, with two regions becoming nonsignificant and one additional region being identified. The magnitude of the LOD score was highly correlated with the heritability of each phenotype as well as the proportion of phenotypic variance due to that locus. There were no false-positive linkage results using the longitudinal data and three false-positive findings using the cross-sectional data. The three false positive results appear to be due to the kurtosis in the trait distribution, even after removing extreme outliers. Our analyses demonstrated that the use of simple longitudinal phenotypes was a powerful means to detect genes of major to moderate effect on trait variability. In only one instance was the power and heritability of the trait increased by using data from one examination. Power to detect linkage can be improved by identifying the most heritable phenotype, ensuring normality of the trait distribution and maximizing the information utilized through novel longitudinal designs for genetic analysis.

Cardiovascular Diseases↗

Assessing the variability in GeneChip data.

INTRODUCTION: Oligonucleotide and cDNA microarray experiments are now common practice in biological science research. The goal of these experiments is generally to gain clues about the functions of genes by measuring how their expression levels rise and fall in response to changing experimental conditions. Measures of gene expression are affected, however, by a variety of factors. This paper introduces statistical methods to assess the variability of Affymetrix GeneChip data due to randomness. METHODS: The variation of Affymetrix's GeneChip signal data are quantified at both chip level and individual gene level, respectively, by the agreement study method and variance components method. Three agreement measurement methods are introduced to assess the variability among chips. Variation sources for gene expression data are decomposed into four categories: systematic experiment variation, treatment effect, biological variation, and chip variation. The focus of this paper is on evaluating and comparing the last two kinds of variations. RESULTS: Measurement of agreement and variance components methods were applied to an experimental data, and the calculation and interpretation were exemplified. The variability between biological samples were shown to exist and were assessed at both the chip level and individual gene level. Using the variance components method, it was found that the biological and chip variation are roughly comparable. The Statistical Analysis System (SAS) program for doing the agreement studies can be obtained from the correspondence author.

Algorithms↗