Search PubMed⌕ Search

Biomedical subjects

Jianjun Chen

Publications and source records attributed to Jianjun Chen.

42 records · Page 3Linked to original sources

Identifying novel transcripts and novel genes in the human genome by using novel SAGE tags.

The number of genes in the human genome is still a controversial issue. Whereas most of the genes in the human genome are said to have been physically or computationally identified, many short cDNA sequences identified as tags by use of serial analysis of gene expression (SAGE) do not match these genes. By performing experimental verification of more than 1,000 SAGE tags and analyzing 4,285,923 SAGE tags of human origin in the current SAGE database, we examined the nature of the unmatched SAGE tags. Our study shows that most of the unmatched SAGE tags are truly novel SAGE tags that originated from novel transcripts not yet identified in the human genome, including alternatively spliced transcripts from known genes and potential novel genes. Our study indicates that by using novel SAGE tags as probes, we should be able to identify efficiently many novel transcripts/novel genes in the human genome that are difficult to identify by conventional methods.

Base Pair Mismatch↗

Molecular portraits of B cell lineage commitment.

In an attempt to characterize early B cell development including the commitment of progenitor cells to the B cell lineage, we generated and compared genomewide gene expression profiles of human hematopoietic stem cells (HSCs) and pre-B cells (PBCs) by using serial analysis of gene expression. From more than 100,000 serial analysis of gene expression tags collected from human CD34(+) HSCs and CD10(+) CD19(+) PBCs, 42,399 unique transcripts were identified in HSCs but only 16,786 in PBCs, suggesting that more than 60% of transcripts expressed in HSCs were silenced during or after commitment to the B cell lineage. On the other hand, mRNAs of pre-B cell receptor (pre-BCR)-associated genes are virtually missing in HSCs but account for more than 10% of the transcriptome of PBCs, which also show increased expression of apoptosis-related genes. Both concentration of the transcriptional repertoire on pre-BCR-related genes together with marked up-regulation of apoptosis mediators in PBC might reflect selection for the expression of a functional pre-BCR within the bone marrow. Besides known regulator genes of early B cell development such as PAX5, E2A, and EBF, the most abundantly expressed genes in PBCs include ATM, PDGFRA, SIAH1, PIM2, C/EBPB, WNT16, and TCL1, the role of which has not been established yet in early B cell development.

Antigens, CD↗

Oligo(dT) primer generates a high frequency of truncated cDNAs through internal poly(A) priming during reverse transcription.

We have analyzed a systematic flaw in the current system of gene identification: the oligo(dT) primer widely used for cDNA synthesis generates a high frequency of truncated cDNAs through internal poly(A) priming. Such truncated cDNAs may contribute to 12% of the expressed sequence tags in the current dbEST database. By using a synthetic transcript and real mRNA templates as models, we characterized the patterns of internal poly(A) priming by oligo(dT) primer. We further demonstrated that the internal poly(A) priming can be effectively diminished by replacing the oligo(dT) primer with a set of anchored oligo(dT) primers for reverse transcription. Our study indicates that cDNAs designed for genomewide gene identification should be synthesized by use of the anchored oligo(dT) primers, rather than the oligo(dT) primers, to diminish the generation of truncated cDNAs caused by internal poly(A) priming.

Animals↗

Genomic DNA breakpoints in AML1/RUNX1 and ETO cluster with topoisomerase II DNA cleavage and DNase I hypersensitive sites in t(8;21) leukemia.

The translocation t(8;21)(q22;q22) is one of the most frequent chromosome translocations in acute myeloid leukemia (AML). AML1/RUNX1 at 21q22 is involved in t(8;21), t(3;21), and t(16;21) in de novo and therapy-related AML and myelodysplastic syndrome as well as in t(12;21) in childhood B cell acute lymphoblastic leukemia. Although DNA breakpoints in AML1 and ETO (at 8q22) cluster in a few introns, the mechanisms of DNA recombination resulting in t(8;21) are unknown. The correlation of specific chromatin structural elements, i.e., topoisomerase II (topo II) DNA cleavage sites, DNase I hypersensitive sites, and scaffold-associated regions, which have been implicated in chromosome recombination with genomic DNA breakpoints in AML1 and ETO in t(8;21) is unknown. The breakpoints in AML1 and ETO were clustered in the Kasumi 1 cell line and in 31 leukemia patients with t(8;21); all except one had de novo AML. Sequencing of the breakpoint junctions revealed no common DNA motif; however, deletions, duplications, microhomologies, and nontemplate DNA were found. Ten in vivo topo II DNA cleavage sites were mapped in AML1, including three in intron 5 and seven in intron 7a, and two were in intron 1b of ETO. All strong topo II sites colocalized with DNase I hypersensitive sites and thus represent open chromatin regions. These sites correlated with genomic DNA breakpoints in both AML1 and ETO, thus implicating them in the de novo 8;21 translocation.

Adult↗

High-throughput GLGI procedure for converting a large number of serial analysis of gene expression tag sequences into 3' complementary DNAs.

Serial analysis of gene expression (SAGE) is a powerful technique for genome-wide analysis of gene expression. However, two-thirds of SAGE tags cannot be used directly for gene identification for two reasons. First, many SAGE tags match several known expressed sequences, owing to the short length of SAGE tag sequences. Second, many SAGE tags do not match any known expressed sequences, presumably because the sequences corresponding to these SAGE tags have not been identified. These two problems can be solved by extension of the SAGE tags into 3' complementary DNAs (cDNAs) by use of the GLGI technique (generation of longer cDNA fragments from SAGE tags for gene identification). We have improved the original GLGI technique into a high-throughput procedure for simultaneous conversion of a large number of SAGE tags into corresponding 3' cDNAs. The whole process is simple, rapid, low-cost, and highly efficient, as shown by our use of this procedure for analyzing hundreds of SAGE tags. In addition to identifying the correct gene for SAGE tags with multiple matches, GLGI can be used for large-scale identification of novel genes by converting novel SAGE tags into 3' cDNAs. Applying this high-throughput procedure should accelerate the rate of gene identification significantly in the human and other eukaryotic genomes.

Animals↗

Correct identification of genes from serial analysis of gene expression tag sequences.

SAGE (serial analysis of gene expression) is a remarkable technique for genome-wide analysis of gene expression. It is crucial to understand the extent to which SAGE can accurately indicate a gene or expressed sequence tag (EST) with a single tag. We analyzed the effect of the size of SAGE tag on gene identification. Our observation indicates that SAGE tags are in general not long enough to achieve the degree of uniqueness of identification originally envisaged. Our observations also indicate that the limitation of using SAGE tag to identify a gene can be overcome by converting SAGE tags into longer 3' EST sequences with the generation of longer cDNA fragments from SAGE tages for gene identification (GLGI) method.

Expressed Sequence Tags↗