Search PubMed⌕ Search

Biomedical subjects

San Ming Wang

Publications and source records attributed to San Ming Wang.

At least 19 recordsLinked to original sources

Understanding SAGE data.

Serial analysis of gene expression (SAGE) is a method for identifying and quantifying transcripts from eukaryotic genomes. Since its invention, SAGE has been widely applied to analyzing gene expression in many biological and medical studies. Vast amounts of SAGE data have been collected and more than a thousand SAGE-related studies have been published since the mid-1990s. The principle of SAGE has been developed to address specific issues such as determination of normal gene structure and identification of abnormal genome structural changes. This review focuses on the general features of SAGE data, including the specificity of SAGE tags with respect to their original transcripts, the quantitative nature of SAGE data for differentially expressed genes, the reproducibility, the comparability of SAGE with microarray and the future potential of SAGE. Understanding these basic features should aid the proper interpretation of SAGE data to address biological and medical questions.

Data Interpretation, Statistical↗

Pan-genome isolation of low abundance transcripts using SAGE tag.

The SAGE (serial analysis of gene expression) method is sensitive at detecting the lower abundance transcripts. More than a third of human SAGE tags identified are novel representing the low abundance unknown transcripts. Using the GLGI method (generation of longer 3' EST from SAGE tag for gene identification), we converted 1009 low-copy, human X chromosome-specific SAGE tags into 10210 3' ESTs. We identified 3418 unique 3' ESTs, 46% of which are novel and originated from the lower abundance transcripts. However, nearly all 3' ESTs were mapped to various regions across the genome but not X chromosome. Detailed analysis indicates that those 3' ESTs were isolated by SAGE tag mis-priming to the non-parent transcripts. Replacing SAGE tags with non-transcribed genomic DNA tags resulted in poor amplification, indicating that the sequence similarity between different transcripts contributed to the amplification. Our study shows the prevalence of novel low abundance transcripts that can be isolated efficiently through SAGE tags mis-priming.

Base Pair Mismatch↗

SAGE detects microRNA precursors.

BACKGROUND: MicroRNAs (miRNAs) have been shown to play important roles in regulating gene expression. Since miRNAs are often evolutionarily conserved and their precursors can be folded into stem-loop hairpins, many miRNAs have been predicted. Yet experimental confirmation is difficult since miRNA expression is often specific to particular tissues and developmental stages. RESULTS: Analysis of 29 human and 230 mouse longSAGE libraries revealed the expression of 22 known and 10 predicted mammalian miRNAs. Most were detected in embryonic tissues. Four SAGE tags detected in human embryonic stem cells specifically match a cluster of four human miRNAs (mir-302a, b, c&d) known to be expressed in embryonic stem cells. LongSAGE data also suggest the existence of a mouse homolog of human and rat mir-493. CONCLUSION: The observation that some orphan longSAGE tags uniquely match miRNA precursors provides information about the expression of some known and predicted miRNAs.

Animals↗

A large quantity of novel human antisense transcripts detected by LongSAGE.

MOTIVATION: Taking advantage of the high sensitivity and specificity of LongSAGE tag for transcript detection and genome mapping, we analyzed the 632 813 unique human LongSAGE tags deposited in public databases to identify novel human antisense transcripts. RESULTS: Our study identified 45 321 tags that match the antisense strand of 9804 known mRNA sequences, 6606 of which contain antisense ESTs and 3198 are mapped only by SAGE tags. Quantitative analysis showed that the detected antisense transcripts are present at levels lower than their counterpart sense transcripts. Experimental results confirmed the presence of antisense transcripts detected by the antisense tags. We also constructed an antisense tag database that can be used to identify the antisense SAGE tags originated from the antisense strand of known mRNA sequences included in the RefSeq database. CONCLUSIONS: Our study highlights the benefits of exploring SAGE data for comprehensive identification of human antisense transcripts and demonstrates the prevalence of antisense transcripts in the human genome.

Base Sequence↗

A novel primate specific gene, CEI, is located in the homeobox gene IRXA2 promoter in Homo sapiens.

The Iroquois (IRX) homeobox gene family consists of six highly conserved transcription factors that are of importance for normal embryonic development. They are organized in two gene clusters in human, one on 5p15.33 and the other one on 16q12.2, respectively, and both the organization and the structure of the genes are highly conserved. An open reading frame coding for an unknown protein is identified in the promoter of IRXA2 on chromosome 5p. This new gene is composed of four exons and it is orientated in a head-to-head manner to IRXA2. Only a short 851 bp segment separates the two translation start codons and the two genes may share a bi-directional promoter. This bi-directional promoter is embedded in a large CpG-island, that also continues into both genes. RT-PCR analysis of the new gene reveals two alternative mRNA transcripts and a third mRNA transcript can be predicted from EST clones. The expression profile of the gene analysed in 9 different human tissues reveals that it is expressed in a coordinated fashion with IRXA2, which has led to the name CEI (Coordinated Expression to IRXA2). The CEI protein lacks homology to any known protein or protein domain in public databases, and a putative amino terminal signal peptide suggests the protein is secreted or ER compartment located. The gene is only found in the human and the chimpanzee genome, but not in the mouse or the rat genome, which suggests that CEI is unique for higher primates. As the identified bi-directional promoter not being a relic of an ancient compact genome, CEI may play an important role in the evolution of higher primates in coordination with the IRX genes.

Amino Acid Sequence↗

Gene expression profiles in acute myeloid leukemia with common translocations using SAGE.

Identification of the specific cytogenetic abnormality is one of the critical steps for classification of acute myeloblastic leukemia (AML) which influences the selection of appropriate therapy and provides information about disease prognosis. However at present, the genetic complexity of AML is only partially understood. To obtain a comprehensive, unbiased, quantitative measure, we performed serial analysis of gene expression (SAGE) on CD15(+) myeloid progenitor cells from 22 AML patients who had four of the most common translocations, namely t(8;21), t(15;17), t(9;11), and inv(16). The quantitative data provide clear evidence that the major change in all these translocation-carrying leukemias is a decrease in expression of the majority of transcripts compared with normal CD15(+) cells. From a total of 1,247,535 SAGE tags, we identified 2,604 transcripts whose expression was significantly altered in these leukemias compared with normal myeloid progenitor cells. The gene ontology of the 1,110 transcripts that matched known genes revealed that each translocation had a uniquely altered profile in various functional categories including regulation of transcription, cell cycle, protein synthesis, and apoptosis. Our global analysis of gene expression of common translocations in AML can focus attention on the function of the genes with altered expression for future biological studies as well as highlight genes/pathways for more specifically targeted therapy.

Apoptosis↗

Applying the SAGE technique to study the effects of electromagnetic field on biological systems.

Identification of genes alternatively expressed in electromagnetic field (EMF)-exposed cells could provide direct evidence for biological effects of EMF. As there are a few indications so far for certain genes to be influenced by EMF, genome-wide scans of the transcriptome appear to be necessary. Among the several technologies used for genome-wide gene expression analysis, serial analysis of gene expression (SAGE) is one promising method, which seems particularly applicable for EMF research. This review provides a brief description of the features of gene expression, illustrates the basic principle of SAGE, and discusses the advantages and limitations of SAGE as well as examples of application. This information should help investigators determine if the SAGE technique is an optimal method for evaluating the biological effects of EMF.

Animals↗

Annotating nonspecific SAGE tags with microarray data.

SAGE (serial analysis of gene expression) detects transcripts by extracting short tags from the transcripts. Because of the limited length, many SAGE tags are shared by transcripts from different genes. Relying on sequence information in the general gene expression database has limited power to solve this problem due to the highly heterogeneous nature of the deposited sequences. Considering that the complexity of gene expression at a single tissue level should be much simpler than that in the general expression database, we reasoned that by restricting gene expression to tissue level, the accuracy of gene annotation for the nonspecific SAGE tags should be significantly improved. To test the idea, we developed a tissue-specific SAGE annotation database based on microarray data (). This database contains microarray expression information represented as UniGene clusters for 73 normal human tissues and 18 cancer tissues and cell lines. The nonspecific SAGE tag is first matched to the database by the same tissue type used by both SAGE and microarray analysis; then the multiple UniGene clusters assigned to the nonspecific SAGE tag are searched in the database under the matched tissue type. The UniGene cluster presented solely or at higher expression levels in the database is annotated to represent the specific gene for the nonspecific SAGE tags. The accuracy of gene annotation by this database was largely confirmed by experimental data. Our study shows that microarray data provide a useful source for annotating the nonspecific SAGE tags.

Cell Line↗

2.45 GHz radiofrequency fields alter gene expression in cultured human cells.

The biological effect of radiofrequency (RF) fields remains controversial. We address this issue by examining whether RF fields can cause changes in gene expression. We used the pulsed RF fields at a frequency of 2.45 GHz that is commonly used in telecommunication to expose cultured human HL-60 cells. We used the serial analysis of gene expression (SAGE) method to measure the RF effect on gene expression at the genome level. We observed that 221 genes altered their expression after a 2-h exposure. The number of affected genes increased to 759 after a 6-h exposure. Functional classification of the affected genes reveals that apoptosis-related genes were among the upregulated ones and the cell cycle genes among the downregulated ones. We observed no significant increase in the expression of heat shock genes. These results indicate that the RF fields at 2.45 GHz can alter gene expression in cultured human cells through non-thermal mechanism.

Dose-Response Relationship, Radiation↗

Stage-dependent gene expression profiles during natural killer cell development.

Natural killer (NK) cells develop from hematopoietic stem cells (HSCs) in the bone marrow. To understand the molecular regulation of NK cell development, serial analysis of gene expression (SAGE) was applied to HSCs, NK precursor (pNK) cells, and mature NK cells (mNK) cultured without or with OP9 stromal cells. From 170,464 total individual tags from four SAGE libraries, 35,385 unique genes were identified. A set of genes was expressed in a stage-specific manner: 15 genes in HSCs, 30 genes in pNK cells, and 27 genes in mNK cells. Among them, lipoprotein lipase induced NK cell maturation and cytotoxic activity. Identification of genome-wide profiles of gene expression in different stages of NK cell development affords us a fundamental basis for defining the molecular network during NK cell development.

Animals↗

Interpreting expression profiles of cancers by genome-wide survey of breadth of expression in normal tissues.

A critical and difficult part of studying cancer with DNA microarrays is data interpretation. Besides the need for data analysis algorithms, integration of additional information about genes might be useful. We performed genome-wide expression profiling of 36 types of normal human tissues and identified 2503 tissue-specific genes. We then systematically studied the expression of these genes in cancers by reanalyzing a large collection of published DNA microarray datasets. We observed that the expression level of liver-specific genes in hepatocellular carcinoma (HCC) correlates with the clinically defined degree of tumor differentiation. Through unsupervised clustering of tissue-specific genes differentially expressed in tumors, we extracted expression patterns that are characteristic of individual cell types, uncovering differences in cell lineage among tumor subtypes. We were able to detect the expression signature of hepatocytes in HCC, neuron cells in medulloblastoma, glia cells in glioma, basal and luminal epithelial cells in breast tumors, and various cell types in lung cancer samples. We also demonstrated that tissue-specific expression signatures are useful in locating the origin of metastatic tumors. Our study shows that integration of each gene's breadth of expression (BOE) in normal tissues is important for biological interpretation of the expression profiles of cancers in terms of tumor differentiation, cell lineage, and metastasis.

Algorithms↗

Serial analysis of gene expression study of a hybrid rice strain (LYP9) and its parental cultivars.

Using the serial analysis of gene expression technique, we surveyed transcriptomes of three major tissues (panicles, leaves, and roots) of a super-hybrid rice (Oryza sativa) strain, LYP9, in comparison to its parental cultivars, 93-11 (indica) and PA64s (japonica). We acquired 465,679 tags from the serial analysis of gene expression libraries, which were consolidated into 68,483 unique tags. Focusing our initial functional analyses on a subset of the data that are supported by full-length cDNAs and the tags (genes) differentially expressed in the hybrid at a significant level (P<0.01), we identified 595 up-regulated (22 tags in panicles, 228 in leaves, and 345 in roots) and 25 down-regulated (seven tags in panicles, 15 in leaves, and three in roots) in LYP9. Most of the tag-identified and up-regulated genes were found related to enhancing carbon- and nitrogen-assimilation, including photosynthesis in leaves, nitrogen uptake in roots, and rapid growth in both roots and panicles. Among the down-regulated genes in LYP9, there is an essential enzyme in photorespiration, alanine:glyoxylate aminotransferase 1. Our study adds a new set of data crucial for the understanding of molecular mechanisms of heterosis and gene regulation networks of the cultivated rice.

Chromosome Mapping↗

Detecting novel low-abundant transcripts in Drosophila.

Increasing evidence suggests that low-abundant transcripts may play fundamental roles in biological processes. In an attempt to estimate the prevalence of low-abundant transcripts in eukaryotic genomes, we performed a transcriptome analysis in Drosophila using the SAGE technique. We collected 244,313 SAGE tags from transcripts expressed in Drosophila embryonic, larval, pupae, adult, and testicular tissue. From these SAGE tags, we identified 40,823 unique SAGE tags. Our analysis showed that 55% of the 40,823 unique SAGE tags are novel without matches in currently known Drosophila transcripts, and most of the novel SAGE tags have low copy numbers. Further analysis indicated that these novel SAGE tags represent novel low-abundant transcripts expressed from loci outside of currently annotated exons including the intergenic and intronic regions, and antisense of the currently annotated exons in the Drosophila genome. Our study reveals the presence of a significant number of novel low-abundant transcripts in Drosophila, and highlights the need to isolate these novel low-abundant transcripts for further biological studies.

Animals↗

Generation of longer 3' cDNA fragments from massively parallel signature sequencing tags.

Massively Parallel Signature Sequencing (MPSS) is a powerful technique for genome-wide gene expression analysis, which, similar to SAGE, relies on the production of short tags proximal to the 3'end of transcripts. A single MPSS experiment can generate over 10(7) tags, providing a 10-fold coverage of the transcripts expressed in a human cell. A significant fraction of MPSS tags cannot be assigned to known transcripts (orphan tags) and are likely to be derived from transcripts expressed at very low levels (approximately 1 copy per cell). In order to explore the potential of MPSS for the characterization of the human transcriptome, we have adapted the GLGI protocol (Generation of Longer cDNA fragments from SAGE tags for Gene Identification) to convert MPSS tags into their corresponding 3' cDNA fragments. GLGI-MPSS was applied to 83 orphan tags and 41 cDNA fragments were obtained. The analysis of these 41 fragments allowed the identification of novel transcripts, alternative tags generated from polymorphic and alternatively spliced transcripts, as well as the detection of artefactual MPSS tags. A systematic large-scale analysis of the genome by MPSS, in combination with the use of GLGI-MPSS protocol, will certainly provide a complementary approach to generate the complete catalog of human transcripts.

Alternative Splicing↗

SAGE is far more sensitive than EST for detecting low-abundance transcripts.

BACKGROUND: Isolation of low-abundance transcripts expressed in a genome remains a serious challenge in transcriptome studies. The sensitivity of the methods used for analysis has a direct impact on the efficiency of the detection. We compared the EST method and the SAGE method to determine which one is more sensitive and to what extent the sensitivity is great for the detection of low-abundance transcripts. RESULTS: Using the same low-abundance transcripts detected by both methods as the targeted sequences, we observed that the SAGE method is 26 times more sensitive than the EST method for the detection of low-abundance transcripts. CONCLUSIONS: The SAGE method is more efficient than the EST method in detecting the low-abundance transcripts.

3' Untranslated Regions↗

Silencing of B cell receptor signals in human naive B cells.

To identify changes in the regulation of B cell receptor (BCR) signals during the development of human B cells, we generated genome-wide gene expression profiles using the serial analysis of gene expression (SAGE) technique for CD34(+) hematopoietic stem cells (HSCs), pre-B cells, naive, germinal center (GC), and memory B cells. Comparing these SAGE profiles, genes encoding positive regulators of BCR signaling were expressed at consistently lower levels in naive B cells than in all other B cell subsets. Conversely, a large group of inhibitory signaling molecules, mostly belonging to the immunoglobulin superfamily (IgSF), were specifically or predominantly expressed in naive B cells. The quantitative differences observed by SAGE were corroborated by semiquantitative reverse transcription-polymerase chain reaction (RT-PCR) and flow cytometry. In a functional assay, we show that down-regulation of inhibitory IgSF receptors and increased responsiveness to BCR stimulation in memory as compared with naive B cells at least partly results from interleukin (IL)-4 receptor signaling. Conversely, activation or impairment of the inhibitory IgSF receptor LIRB1 affected BCR-dependent Ca(2+) mobilization only in naive but not memory B cells. Thus, LIRB1 and IL-4 may represent components of two nonoverlapping gene expression programs in naive and memory B cells, respectively: in naive B cells, a large group of inhibitory IgSF receptors can elevate the BCR signaling threshold to prevent these cells from premature activation and clonal expansion before GC-dependent affinity maturation. In memory B cells, facilitated responsiveness upon reencounter of the immunizing antigen may result from amplification of BCR signals at virtually all levels of signal transduction.

Antigens, CD34↗