Search PubMed⌕ Search

Biomedical subjects

Philipp Kapranov

Publications and source records attributed to Philipp Kapranov.

10 recordsLinked to original sources

Biological function of unannotated transcription during the early development of Drosophila melanogaster.

Many animal and plant genomes are transcribed much more extensively than current annotations predict. However, the biological function of these unannotated transcribed regions is largely unknown. Approximately 7% and 23% of the detected transcribed nucleotides during D. melanogaster embryogenesis map to unannotated intergenic and intronic regions, respectively. Based on computational analysis of coordinated transcription, we conservatively estimate that 29% of all unannotated transcribed sequences function as missed or alternative exons of well-characterized protein-coding genes. We estimate that 15.6% of intergenic transcribed regions function as missed or alternative transcription start sites (TSS) used by 11.4% of the expressed protein-coding genes. Identification of P element mutations within or near newly identified 5' exons provides a strategy for mapping previously uncharacterized mutations to their respective genes. Collectively, these data indicate that at least 85% of the fly genome is transcribed and processed into mature transcripts representing at least 30% of the fly genome.

Amino Acid Sequence↗

CD127 expression inversely correlates with FoxP3 and suppressive function of human CD4+ T reg cells.

Regulatory T (T reg) cells are critical regulators of immune tolerance. Most T reg cells are defined based on expression of CD4, CD25, and the transcription factor, FoxP3. However, these markers have proven problematic for uniquely defining this specialized T cell subset in humans. We found that the IL-7 receptor (CD127) is down-regulated on a subset of CD4(+) T cells in peripheral blood. We demonstrate that the majority of these cells are FoxP3(+), including those that express low levels or no CD25. A combination of CD4, CD25, and CD127 resulted in a highly purified population of T reg cells accounting for significantly more cells that previously identified based on other cell surface markers. These cells were highly suppressive in functional suppressor assays. In fact, cells separated based solely on CD4 and CD127 expression were anergic and, although representing at least three times the number of cells (including both CD25(+)CD4(+) and CD25(-)CD4(+) T cell subsets), were as suppressive as the "classic" CD4(+)CD25(hi) T reg cell subset. Finally, we show that CD127 can be used to quantitate T reg cell subsets in individuals with type 1 diabetes supporting the use of CD127 as a biomarker for human T reg cells.

Adolescent↗

Microarray-based DNA methylation profiling: technology and applications.

This work is dedicated to the development of a technology for unbiased, high-throughput DNA methylation profiling of large genomic regions. In this method, unmethylated and methylated DNA fractions are enriched using a series of treatments with methylation sensitive restriction enzymes, and interrogated on microarrays. We have investigated various aspects of the technology including its replicability, informativeness, sensitivity and optimal PCR conditions using microarrays containing oligonucleotides representing 100 kb of genomic DNA derived from the chromosome 22 COMT region in addition to 12 192 element CpG island microarrays. Several new aspects of methylation profiling are provided, including the parallel identification of confounding effects of DNA sequence variation, the description of the principles of microarray design for epigenomic studies and the optimal choice of methylation sensitive restriction enzymes. We also demonstrate the advantages of using the unmethylated DNA fraction versus the methylated one, which substantially improve the chances of detecting DNA methylation differences. We applied this methodology for fine-mapping of methylation patterns of chromosomes 21 and 22 in eight individuals using tiling microarrays consisting of over 340 000 oligonucleotide probe pairs. The principles developed in this work will help to make epigenetic profiling of the entire human genome a routine procedure.

Chromosome Mapping↗

Temporal profile of replication of human chromosomes.

Chromosomes in human cancer cells are expected to initiate replication from predictably localized origins, firing reproducibly at discrete times in S phase. Replication products obtained from HeLa cells at different stages of S phase were hybridized to cDNA and genome tiling oligonucleotide microarrays to determine the temporal profile of replication of human chromosomes on a genome-wide scale. About 1,000 genes and chromosomal segments were identified as sites containing efficient origins that fire reproducibly. Early replication was correlated with high gene density. An acute transition of gene density from early to late replicating areas suggests that discrete chromatin states dictate early versus late replication. Surprisingly, at least 60% of the interrogated chromosomal segments replicate equally in all quarters of S phase, suggesting that large stretches of chromosomes are replicated by inefficient, variably located and asynchronous origins and forks, producing a pan-S phase pattern of replication. Thus, at least for aneuploid cancer cells, a typical discrete time of replication in S phase is not seen for large segments of the chromosomes.

Chromosomes, Human↗

Transcriptional maps of 10 human chromosomes at 5-nucleotide resolution.

Sites of transcription of polyadenylated and nonpolyadenylated RNAs for 10 human chromosomes were mapped at 5-base pair resolution in eight cell lines. Unannotated, nonpolyadenylated transcripts comprise the major proportion of the transcriptional output of the human genome. Of all transcribed sequences, 19.4, 43.7, and 36.9% were observed to be polyadenylated, nonpolyadenylated, and bimorphic, respectively. Half of all transcribed sequences are found only in the nucleus and for the most part are unannotated. Overall, the transcribed portions of the human genome are predominantly composed of interlaced networks of both poly A+ and poly A- annotated transcripts and unannotated transcripts of unknown function. This organization has important implications for interpreting genotype-phenotype associations, regulation of gene expression, and the definition of a gene.

Cell Line↗

Examples of the complex architecture of the human transcriptome revealed by RACE and high-density tiling arrays.

Recently, we mapped the sites of transcription across approximately 30% of the human genome and elucidated the structures of several hundred novel transcripts. In this report, we describe a novel combination of techniques including the rapid amplification of cDNA ends (RACE) and tiling array technologies that was used to further characterize transcripts in the human transcriptome. This technical approach allows for several important pieces of information to be gathered about each array-detected transcribed region, including strand of origin, start and termination positions, and the exonic structures of spliced and unspliced coding and noncoding RNAs. In this report, the structures of transcripts from 14 transcribed loci, representing both known genes and unannotated transcripts taken from the several hundred randomly selected unannotated transcripts described in our previous work are represented as examples of the complex organization of the human transcriptome. As a consequence of this complexity, it is not unusual that a single base pair can be part of an intricate network of multiple isoforms of overlapping sense and antisense transcripts, the majority of which are unannotated. Some of these transcripts follow the canonical splicing rules, whereas others combine the exons of different genes or represent other types of noncanonical transcripts. These results have important implications concerning the correlation of genotypes to phenotypes, the regulation of complex interlaced transcriptional patterns, and the definition of a gene.

Cell Line↗

Unbiased mapping of transcription factor binding sites along human chromosomes 21 and 22 points to widespread regulation of noncoding RNAs.

Using high-density oligonucleotide arrays representing essentially all nonrepetitive sequences on human chromosomes 21 and 22, we map the binding sites in vivo for three DNA binding transcription factors, Sp1, cMyc, and p53, in an unbiased manner. This mapping reveals an unexpectedly large number of transcription factor binding site (TFBS) regions, with a minimal estimate of 12,000 for Sp1, 25,000 for cMyc, and 1600 for p53 when extrapolated to the full genome. Only 22% of these TFBS regions are located at the 5' termini of protein-coding genes while 36% lie within or immediately 3' to well-characterized genes and are significantly correlated with noncoding RNAs. A significant number of these noncoding RNAs are regulated in response to retinoic acid, and overlapping pairs of protein-coding and noncoding RNAs are often coregulated. Thus, the human genome contains roughly comparable numbers of protein-coding and noncoding genes that are bound by common transcription factors and regulated by common environmental signals.

Amino Acid Motifs↗

Novel RNAs identified from an in-depth analysis of the transcriptome of human chromosomes 21 and 22.

In this report, we have achieved a richer view of the transcriptome for Chromosomes 21 and 22 by using high-density oligonucleotide arrays on cytosolic poly(A)(+) RNA. Conservatively, only 31.4% of the observed transcribed nucleotides correspond to well-annotated genes, whereas an additional 4.8% and 14.7% correspond to mRNAs and ESTs, respectively. Approximately 85% of the known exons were detected, and up to 21% of known genes have only a single isoform based on exon-skipping alternative expression. Overall, the expression of the well-characterized exons falls predominately into two categories, uniquely or ubiquitously expressed with an identifiable proportion of antisense transcripts. The remaining observed transcription (49.0%) was outside of any known annotation. These novel transcripts appear to be more cell-line-specific and have lower and less variation in expression than the well-characterized genes. Novel transcripts were further characterized based on their distance to annotations, transcript size, coding capacity, and identification as antisense to intronic sequences. By RT-PCR, 126 novel transcripts were independently verified, resulting in a 65% verification rate. These observations strongly support the argument for a re-evaluation of the total number of human genes and an alternative term for "gene" to encompass these growing, novel classes of RNA transcripts in the human genome.

Cell Line↗

Beyond expression profiling: next generation uses of high density oligonucleotide arrays.

In the past decade, microarray technology has become a major tool for high-throughput comprehensive analysis of gene expression, genotyping and resequencing applications. Currently, the most widely employed application of high-density oligonucleotide arrays (HDOAs) involves monitoring changes in gene expression. This application has been carried out in a variety of organisms ranging from Escherichia coli to humans. The recent near completion of the human and mouse genome sequences, however, as well as the genomes of other model experimental species, has allowed for novel applications of HDOAs, such as: the discovery of novel transcripts, mapping functionally important genomic regions and identifying functional domains in RNA molecules. Integrating all this information will provide novel global views of the locations of RNA transcription, DNA replication and the protein nucleic acid interactions that regulate these processes.

DNA Replication↗

Large-scale transcriptional activity in chromosomes 21 and 22.

The sequences of the human chromosomes 21 and 22 indicate that there are approximately 770 well-characterized and predicted genes. In this study, empirically derived maps identifying active areas of RNA transcription on these chromosomes have been constructed with the use of cytosolic polyadenylated RNA obtained from 11 human cell lines. Oligonucleotide arrays containing probes spaced on average every 35 base pairs along these chromosomes were used. When compared with the sequence annotations available for these chromosomes, it is noted that as much as an order of magnitude more of the genomic sequence is transcribed than accounted for by the predicted and characterized exons.

Cell Line↗