Search PubMed⌕ Search

Biomedical subjects

Paul Pavlidis

Publications and source records attributed to Paul Pavlidis.

11 recordsLinked to original sources

Sex genes for genomic analysis in human brain: internal controls for comparison of probe level data extraction.

BACKGROUND: Genomic studies of complex tissues pose unique analytical challenges for assessment of data quality, performance of statistical methods used for data extraction, and detection of differentially expressed genes. Ideally, to assess the accuracy of gene expression analysis methods, one needs a set of genes which are known to be differentially expressed in the samples and which can be used as a "gold standard". We introduce the idea of using sex-chromosome genes as an alternative to spiked-in control genes or simulations for assessment of microarray data and analysis methods. RESULTS: Expression of sex-chromosome genes were used as true internal biological controls to compare alternate probe-level data extraction algorithms (Microarray Suite 5.0 [MAS5.0], Model Based Expression Index [MBEI] and Robust Multi-array Average [RMA]), to assess microarray data quality and to establish some statistical guidelines for analyzing large-scale gene expression. These approaches were implemented on a large new dataset of human brain samples. RMA-generated gene expression values were markedly less variable and more reliable than MAS5.0 and MBEI-derived values. A statistical technique controlling the false discovery rate was applied to adjust for multiple testing, as an alternative to the Bonferroni method, and showed no evidence of false negative results. Fourteen probesets, representing nine Y- and two X-chromosome linked genes, displayed significant sex differences in brain prefrontal cortex gene expression. CONCLUSION: In this study, we have demonstrated the use of sex genes as true biological internal controls for genomic analysis of complex tissues, and suggested analytical guidelines for testing alternate oligonucleotide microarray data extraction protocols and for adjusting multiple statistical analysis of differentially expressed genes. Our results also provided evidence for sex differences in gene expression in the brain prefrontal cortex, supporting the notion of a putative direct role of sex-chromosome genes in differentiation and maintenance of sexual dimorphism of the central nervous system. Importantly, these analytical approaches are applicable to all microarray studies that include male and female human or animal subjects.

Algorithms↗

The effect of replication on gene expression microarray experiments.

MOTIVATION: We examine the effect of replication on the detection of apparently differentially expressed genes in gene expression microarray experiments. Our analysis is based on a random sampling approach using real data sets from 16 published studies. We consider both the ability to find genes that meet particular statistical criteria as well as the stability of the results in the face of changing levels of replication. RESULTS: While dependent on the data source, our findings suggest that stable results are typically not obtained until at least five biological replicates have been used. Conversely, for most studies, 10-15 replicates yield results that are quite stable, and there is less improvement in stability as the number of replicates is further increased. Our methods will be of use in evaluating existing data sets and in helping to design new studies.

Animals↗

Hierarchical model of gene regulation by transforming growth factor beta.

Transforming growth factor betas (TGF-betas) regulate key aspects of embryonic development and major human diseases. Although Smad2, Smad3, and extracellular signal-regulated kinase (ERK) mitogen-activated protein kinases (MAPKs) have been proposed as key mediators in TGF-beta signaling, their functional specificities and interactivity in controlling transcriptional programs in different cell types and (patho)physiological contexts are not known. We investigated expression profiles of genes controlled by TGF-beta in fibroblasts with ablations of Smad2, Smad3, and ERK MAPK. Our results suggest that Smad3 is the essential mediator of TGF-beta signaling and directly activates genes encoding regulators of transcription and signal transducers through Smad3/Smad4 DNA-binding motif repeats that are characteristic for immediate-early target genes of TGF-beta but absent in intermediate target genes. In contrast, Smad2 and ERK predominantly transmodulated regulation of both immediate-early and intermediate genes by TGF-beta/Smad3. These results suggest a previously uncharacterized hierarchical model of gene regulation by TGF-beta in which TGF-beta causes direct activation by Smad3 of cascades of regulators of transcription and signaling that are transmodulated by Smad2 and/or ERK.

Animals↗

Inducible enhancement of memory storage and synaptic plasticity in transgenic mice expressing an inhibitor of ATF4 (CREB-2) and C/EBP proteins.

To examine the role of C/EBP-related transcription factors in long-term synaptic plasticity and memory storage, we have used the tetracycline-regulated system and expressed in the forebrain of mice a broad dominant-negative inhibitor of C/EBP (EGFP-AZIP), which preferentially interacts with several inhibiting isoforms of C/EBP. EGFP-AZIP also reduces the expression of ATF4, a distant member of the C/EBP family of transcription factors that is homologous to the Aplysia memory suppressor gene ApCREB-2. Consistent with the removal of inhibitory constraints on transcription, we find an increase in the pattern of gene transcripts in the hippocampus of EGFP-AZIP transgenic mice and both a reversibly enhanced hippocampal-based spatial memory and LTP. These results suggest that several proteins within the C/EBP family including ATF4 (CREB-2) act to constrain long-term synaptic changes and memory formation. Relief of this inhibition lowers the threshold for hippocampal-dependent long-term synaptic potentiation and memory storage in mice.

Activating Transcription Factor 4↗

Classification of clear-cell sarcoma as a subtype of melanoma by genomic profiling.

PURPOSE: To develop a genome-based classification scheme for clear-cell sarcoma (CCS), also known as melanoma of soft parts (MSP), which would have implications for diagnosis and treatment. This tumor displays characteristic features of soft tissue sarcoma (STS), including deep soft tissue primary location and a characteristic translocation, t(12;22)(q13;q12), involving EWS and ATF1 genes. CCS/MSP also has typical melanoma features, including immunoreactivity for S100 and HMB45, pigmentation, MITF-M expression, and a propensity for regional lymph node metastases. MATERIALS AND METHODS: RNA samples from 21 cell lines and 60 pathologically confirmed cases of STS, melanoma, and CCS/MSP were examined using the U95A GeneChip (Affymetrix, Santa Clara, CA). Hierarchical cluster analysis, principal component analysis, and support vector machine (SVM) analysis exploited genomic correlations within the data to classify CCS/MSP. RESULTS: Unsupervised analyses demonstrated a clear distinction between STS and melanoma and, furthermore, showed that CCS/MSP cluster with the melanomas as a distinct group. A supervised SVM learning approach further validated this finding and provided a user-independent approach to diagnosis. Genes of interest that discriminate CCS/MSP included those encoding melanocyte differentiation antigens, MITF, SOX10, ERBB3, and FGFR1. CONCLUSION: Gene expression profiles support the classification of CCS/MSP as a distinct genomic subtype of melanoma. Analysis of these gene profiles using the SVM may be an important diagnostic tool. Genomic analysis identified potential targets for the development of therapeutic strategies in the treatment of this disease.

Algorithms↗

Matrix2png: a utility for visualizing matrix data.

UNLABELLED: We describe a simple software tool, 'matrix2png', for creating color images of matrix data. Originally designed with the display of microarray data sets in mind, it is a general tool that can be used to make simple visualizations of matrices for use in figures, web pages, slide presentations and the like. It can also be used to generate images 'on the fly' in web applications. Both continuous-valued and discrete-valued (categorical) data sets can be displayed. Many options are available to the user, including the colors used, the display of row and column labels, and scale bars. In this note we describe some of matrix2png's features and describe some places it has been useful in the authors' work. AVAILABILITY: A simple web interface is available, and Unix binaries are available from http://microarray.cpmc.columbia.edu/matrix2png. Source code is available on request.

Color↗

Classification and subtype prediction of adult soft tissue sarcoma by functional genomics.

Adult soft tissue sarcomas are a heterogeneous group of tumors, including well-described subtypes by histological and genotypic criteria, and pleomorphic tumors typically characterized by non-recurrent genetic aberrations and karyotypic heterogeneity. The latter pose a diagnostic challenge, even to experienced pathologists. We proposed that gene expression profiling in soft tissue sarcoma would identify a genomic-based classification scheme that is useful in diagnosis. RNA samples from 51 pathologically confirmed cases, representing nine different histological subtypes of adult soft tissue sarcoma, were examined using the Affymetrix U95A GeneChip. Statistical tests were performed on experimental groups identified by cluster analysis, to find discriminating genes that could subsequently be applied in a support vector machine algorithm. Synovial sarcomas, round-cell/myxoid liposarcomas, clear-cell sarcomas and gastrointestinal stromal tumors displayed remarkably distinct and homogenous gene expression profiles. Pleomorphic tumors were heterogeneous. Notably, a subset of malignant fibrous histiocytomas, a controversialhistological subtype, was identified as a distinct genomic group. The support vector machine algorithm supported a genomic basis for diagnosis, with both high sensitivity and specificity. In conclusion, we showed gene expression profiling to be useful in classification and diagnosis, providing insights into pathogenesis and pointing to potential new therapeutic targets of soft tissue sarcoma.

Adult↗

Cutting edge: STAT6 serves as a positive and negative regulator of gene expression in IL-4-stimulated B lymphocytes.

STAT6 plays an important role in IL-4-mediated B cell activation and differentiation. To identify primary and secondary target genes of STAT6, gene expression profiles of IL-4-stimulated B cells from STAT6+/+ vs STAT6-/- mice were compared. Statistical analysis revealed that 106 distinct probe sets including 70 known genes were differentially expressed between the 2 genotypes. These genes include transcription factors, kinases, and other enzymes, cell surface receptors, and Ig H chains. Surprisingly, although 31 genes were expressed at higher levels in STAT6+/+ B cells, 39 genes were expressed at higher abundance in STAT6-/- B cells. This result implies both positive and negative regulatory functions of STAT6 in IL-4-mediated gene expression. Furthermore, IL-4 induces expression of the transcription factor Krox20, which is required for maximal IL-4-induced transcription.

Animals↗

Learning gene functional classifications from multiple data types.

In our attempts to understand cellular function at the molecular level, we must be able to synthesize information from disparate types of genomic data. We consider the problem of inferring gene functional classifications from a heterogeneous data set consisting of DNA microarray expression measurements and phylogenetic profiles from whole-genome sequence comparisons. We demonstrate the application of the support vector machine (SVM) learning algorithm to this functional inference task. Our results suggest the importance of exploiting prior information about the heterogeneity of the data. In particular, we propose an SVM kernel function that is explicitly heterogeneous. In addition, we describe feature scaling methods for further exploiting prior knowledge of heterogeneity by giving each data type different weights.

Algorithms↗

Exploring gene expression data with class scores.

We address a commonly asked question about gene expression data sets: "What functional classes of genes are most interesting in the data?" In the methods we present, expression data is partitioned into classes based on existing annotation schemes. Each class is then given three separately derived "interest" scores. The first score is based on an assessment of the statistical significance of gene expression changes experienced by members of the class, in the context of the experimental design. The second is based on the co-expression of genes in the class. The third score is based on the learnability of the classification. We show that all three methods reveal significant classes in each of three different gene expression data sets. Many classes are identified by one method but not the others, indicating that the methods are complementary. The classes identified are in many cases of clear relevance to the experiment. Our results suggest that these class scoring methods are useful tools for exploring gene expression data.

Analysis of Variance↗

Differential amplification of gene expression in lens cell lines conditioned to survive peroxide stress.

PURPOSE: The response of lens systems to oxidative stress is confusing. Antioxidative defense systems are not mobilized as expected, and unanticipated defenses appear important. Therefore, mouse lens cell lines conditioned to survive different peroxide stresses have been analyzed to determine their global changes in gene expression. METHODS: The immortal mouse lens epithelial cell line alphaTN4-1 was conditioned to survive 125 microM H2O2 (H cells) or a combination of both 100 microM tertiary butyl hydroperoxide (TBHP) and 125 microM H2O2 (HT cells), by a methodology previously described. The total RNA was isolated from the different cell lines and analyzed with oligonucleotide mouse expression microarrays. Four microarrays were used for each cell line. Microarray results were confirmed by real-time RT-PCR. RESULTS: A new cell line resistant to both 125 microM H2O2 and 100 micro M TBHP was developed, because cells resistant to H2O2 were killed by TBHP. Analysis of classic antioxidative enzyme activities showed little change between cells that survive H2O2 (H) and those that survive H2O2 and TBHP (HT). Therefore, the global change in gene expression in these cell lines was determined with gene expression microarrays. The fluorescent signal changes of the genes within the three cell lines, H, HT, and control (C), were analyzed by statistical methods including Tukey analysis. It was found that from the 12,422 gene fragments and expressed sequence tags (ESTs) analyzed--based on a one-way ANOVA with a stringent cutoff of one false positive per 1000 genes and correcting for microarray background and noise--approximately 950 (7.6%) genes had a significant change in expression in comparing the C, H, and HT groups. A small group of antioxidative defense genes were found in this population, including catalase, members of the glutathione (GSH)-S-transferase family, NAD(P)H menadione oxidoreductase 1, and the ferritin light chain. The remaining genes are involved in a broad spectrum of other biological systems. In the HT versus H comparison, only a few genes were found that had increased expression in the HT line compared with expression in the H line, including GSH-S-transferase alpha 3 and hephaestin. Many genes that are frequently considered antioxidative defense genes, including most of the GSH peroxidases, unexpectedly showed little change. CONCLUSIONS: An unusual and generally unexpected small group of antioxidative defense genes appear to have increased expression in response to H2O2 stress. Cell lines resistant to H2O2 do not appear to survive challenge with another type of peroxide, TBHP, a lipid peroxide prototype. However, acquisition of TBHP resistance by H cells was found to be accompanied by significantly amplified expression of only a few additional antioxidative defense genes. Many of the amplified genes do not appear to be involved with antioxidative systems, reflecting the complexity of the cells' response to oxidative stress.

Animals↗