Search PubMed⌕ Search

Biomedical subjects

Thomas R Gingeras

Publications and source records attributed to Thomas R Gingeras.

At least 19 recordsLinked to original sources

Genome-wide analysis of polymerase III-transcribed Alu elements suggests cell-type-specific enhancer function.

Alu elements are one of the most successful families of transposons in the human genome. A portion of Alu elements is transcribed by RNA Pol III, whereas the remaining ones are part of Pol II transcripts. Because Alu elements are highly repetitive, it has been difficult to identify the Pol III-transcribed elements and quantify their expression levels. In this study, we generated high-resolution, long-genomic-span RAMPAGE data in 155 biosamples all with matching RNA-seq data and built an atlas of 17,249 Pol III-transcribed Alu elements. We further performed an integrative analysis on the ChIP-seq data of 10 histone marks and hundreds of transcription factors, whole-genome bisulfite sequencing data, ChIA-PET data, and functional data in several biosamples, and our results revealed that although the human-specific Alu elements are transcriptionally repressed, the older, expressed Alu elements may be exapted by the human host to function as cell-type-specific enhancers for their nearby protein-coding genes.

Alu Elements↗

Relationships between p63 binding, DNA sequence, transcription activity, and biological function in human cells.

Using tiled microarrays covering the entire human genome, we identify approximately 5800 target sites for p63, a p53 homolog essential for stratified epithelial development. p63 targets are enriched for genes involved in cell adhesion, proliferation, death, and signaling pathways. The quality of the derived DNA sequence motif for p63 targets correlates with binding strength binding in vivo, but only a small minority of motifs in the genome is bound by p63. Conversely, many p63 targets have motif scores expected for random genomic regions. Thus, p63 binding in vivo is highly selective and often requires additional factors beyond the simple protein-DNA interaction. There is a significant, but complex, relationship between p63 target sites and p63-responsive genes, with DeltaNp63 isoforms being linked to transcriptional activation. Many p63 binding regions are evolutionarily conserved and/or associated with sequence motifs for other transcription factors, suggesting that a substantial portion of p63 sites is biologically relevant.

Apoptosis↗

Rank-statistics based enrichment-site prediction algorithm developed for chromatin immunoprecipitation on chip experiments.

BACKGROUND: High density oligonucleotide tiling arrays are an effective and powerful platform for conducting unbiased genome-wide studies. The ab initio probe selection method employed in tiling arrays is unbiased, and thus ensures consistent sampling across coding and non-coding regions of the genome. Tiling arrays are increasingly used in chromatin immunoprecipitation (IP) experiments (ChIP on chip). ChIP on chip facilitates the generation of genome-wide maps of in-vivo interactions between DNA-associated proteins including transcription factors and DNA. Analysis of the hybridization of an immunoprecipitated sample to a tiling array facilitates the identification of ChIP-enriched segments of the genome. These enriched segments are putative targets of antibody assayable regulatory elements. The enrichment response is not ubiquitous across the genome. Typically 5 to 10% of tiled probes manifest some significant enrichment. Depending upon the factor being studied, this response can drop to less than 1%. The detection and assessment of significance for interactions that emanate from non-canonical and/or un-annotated regions of the genome is especially challenging. This is the motivation behind the proposed algorithm. RESULTS: We have proposed a novel rank and replicate statistics-based methodology for identifying and ascribing statistical confidence to regions of ChIP-enrichment. The algorithm is optimized for identification of sites that manifest low levels of enrichment but are true positives, as validated by alternative biochemical experiments. Although the method is described here in the context of ChIP on chip experiments, it can be generalized to any treatment-control experimental design. The results of the algorithm show a high degree of concordance with independent biochemical validation methods. The sensitivity and specificity of the algorithm have been characterized via quantitative PCR and independent computational approaches. CONCLUSION: The algorithm ranks all enrichment sites based on their intra-replicate ranks and inter-replicate rank consistency. Following the ranking, the method allows segmentation of sites based on a meta p-value, a composite array signal enrichment criterion, or a composite of these two measures. The sensitivities obtained subsequent to the segmentation of data using a meta p-value of 10-5, an array signal enrichment of 0.2 and a composite of these two values are 88%, 87% and 95%, respectively.

Algorithms↗

Genome-wide analysis of estrogen receptor binding sites.

The estrogen receptor is the master transcriptional regulator of breast cancer phenotype and the archetype of a molecular therapeutic target. We mapped all estrogen receptor and RNA polymerase II binding sites on a genome-wide scale, identifying the authentic cis binding sites and target genes, in breast cancer cells. Combining this unique resource with gene expression data demonstrates distinct temporal mechanisms of estrogen-mediated gene regulation, particularly in the case of estrogen-suppressed genes. Furthermore, this resource has allowed the identification of cis-regulatory sites in previously unexplored regions of the genome and the cooperating transcription factors underlying estrogen signaling in breast cancer.

Adaptor Proteins, Signal Transducing↗

Biological function of unannotated transcription during the early development of Drosophila melanogaster.

Many animal and plant genomes are transcribed much more extensively than current annotations predict. However, the biological function of these unannotated transcribed regions is largely unknown. Approximately 7% and 23% of the detected transcribed nucleotides during D. melanogaster embryogenesis map to unannotated intergenic and intronic regions, respectively. Based on computational analysis of coordinated transcription, we conservatively estimate that 29% of all unannotated transcribed sequences function as missed or alternative exons of well-characterized protein-coding genes. We estimate that 15.6% of intergenic transcribed regions function as missed or alternative transcription start sites (TSS) used by 11.4% of the expressed protein-coding genes. Identification of P element mutations within or near newly identified 5' exons provides a strategy for mapping previously uncharacterized mutations to their respective genes. Collectively, these data indicate that at least 85% of the fly genome is transcribed and processed into mature transcripts representing at least 30% of the fly genome.

Amino Acid Sequence↗

EGASP: the human ENCODE Genome Annotation Assessment Project.

BACKGROUND: We present the results of EGASP, a community experiment to assess the state-of-the-art in genome annotation within the ENCODE regions, which span 1% of the human genome sequence. The experiment had two major goals: the assessment of the accuracy of computational methods to predict protein coding genes; and the overall assessment of the completeness of the current human genome annotations as represented in the ENCODE regions. For the computational prediction assessment, eighteen groups contributed gene predictions. We evaluated these submissions against each other based on a 'reference set' of annotations generated as part of the GENCODE project. These annotations were not available to the prediction groups prior to the submission deadline, so that their predictions were blind and an external advisory committee could perform a fair assessment. RESULTS: The best methods had at least one gene transcript correctly predicted for close to 70% of the annotated genes. Nevertheless, the multiple transcript accuracy, taking into account alternative splicing, reached only approximately 40% to 50% accuracy. At the coding nucleotide level, the best programs reached an accuracy of 90% in both sensitivity and specificity. Programs relying on mRNA and protein sequences were the most accurate in reproducing the manually curated annotations. Experimental validation shows that only a very small percentage (3.2%) of the selected 221 computationally predicted exons outside of the existing annotation could be verified. CONCLUSION: This is the first such experiment in human DNA, and we have followed the standards established in a similar experiment, GASP1, in Drosophila melanogaster. We believe the results presented here contribute to the value of ongoing large-scale annotation projects and should guide further experimental methods when being scaled up to the entire human genome sequence.

Alternative Splicing↗

CD127 expression inversely correlates with FoxP3 and suppressive function of human CD4+ T reg cells.

Regulatory T (T reg) cells are critical regulators of immune tolerance. Most T reg cells are defined based on expression of CD4, CD25, and the transcription factor, FoxP3. However, these markers have proven problematic for uniquely defining this specialized T cell subset in humans. We found that the IL-7 receptor (CD127) is down-regulated on a subset of CD4(+) T cells in peripheral blood. We demonstrate that the majority of these cells are FoxP3(+), including those that express low levels or no CD25. A combination of CD4, CD25, and CD127 resulted in a highly purified population of T reg cells accounting for significantly more cells that previously identified based on other cell surface markers. These cells were highly suppressive in functional suppressor assays. In fact, cells separated based solely on CD4 and CD127 expression were anergic and, although representing at least three times the number of cells (including both CD25(+)CD4(+) and CD25(-)CD4(+) T cell subsets), were as suppressive as the "classic" CD4(+)CD25(hi) T reg cell subset. Finally, we show that CD127 can be used to quantitate T reg cell subsets in individuals with type 1 diabetes supporting the use of CD127 as a biomarker for human T reg cells.

Adolescent↗

TUF love for "junk" DNA.

The widespread occurrence of noncoding (nc) RNAs--unannotated eukaryotic transcripts with reduced protein coding potential--suggests that they are functionally important. Study of ncRNAs is increasing our understanding of the organization and regulation of genomes.

Animals↗

HIV regulation of the IL-7R: a viral mechanism for enhancing HIV-1 replication in human macrophages in vitro.

We report a novel mechanism, involving up-regulation of the interleukin (IL)-7 cytokine receptor, by which human immunodeficiency virus (HIV) enhances its own production in monocyte-derived macrophages (MDM) in vitro. HIV-1 infection or treatment of MDM cultures with exogenous HIV-1 Tat(86) protein up-regulates the IL-7 receptor (IL-7R) alpha-chain at the levels of steady-state RNA, protein, and functional IL-7R on the cell surface (as measured by ligand-induced receptor signaling). This IL-7R up-regulation is associated with increased amounts of HIV-1 virions in the supernatants of infected MDM cultures treated with exogenous IL-7 cytokine. The overall effect of IL-7 stimulation on HIV replication in MDM culture supernatants is typically in the range of one log and greater. The results are consistent with a model in which HIV infection produces the Tat protein, which in turn up-regulates IL-7R in a paracrine manner. This results in increased IL-7R signaling in response to the IL-7 cytokine, which ultimately promotes early events in HIV replication, including binding/entry and possibly other steps prior to reverse transcription. The results suggest that the effects of IL-7 on HIV replication in MDM should be considered when analyzing and designing clinical trials involving treatment of patients with IL-7 or Tat vaccines.

Cells, Cultured↗

Temporal profile of replication of human chromosomes.

Chromosomes in human cancer cells are expected to initiate replication from predictably localized origins, firing reproducibly at discrete times in S phase. Replication products obtained from HeLa cells at different stages of S phase were hybridized to cDNA and genome tiling oligonucleotide microarrays to determine the temporal profile of replication of human chromosomes on a genome-wide scale. About 1,000 genes and chromosomal segments were identified as sites containing efficient origins that fire reproducibly. Early replication was correlated with high gene density. An acute transition of gene density from early to late replicating areas suggests that discrete chromatin states dictate early versus late replication. Surprisingly, at least 60% of the interrogated chromosomal segments replicate equally in all quarters of S phase, suggesting that large stretches of chromosomes are replicated by inefficient, variably located and asynchronous origins and forks, producing a pan-S phase pattern of replication. Thus, at least for aneuploid cancer cells, a typical discrete time of replication in S phase is not seen for large segments of the chromosomes.

Chromosomes, Human↗

Transcriptional maps of 10 human chromosomes at 5-nucleotide resolution.

Sites of transcription of polyadenylated and nonpolyadenylated RNAs for 10 human chromosomes were mapped at 5-base pair resolution in eight cell lines. Unannotated, nonpolyadenylated transcripts comprise the major proportion of the transcriptional output of the human genome. Of all transcribed sequences, 19.4, 43.7, and 36.9% were observed to be polyadenylated, nonpolyadenylated, and bimorphic, respectively. Half of all transcribed sequences are found only in the nucleus and for the most part are unannotated. Overall, the transcribed portions of the human genome are predominantly composed of interlaced networks of both poly A+ and poly A- annotated transcripts and unannotated transcripts of unknown function. This organization has important implications for interpreting genotype-phenotype associations, regulation of gene expression, and the definition of a gene.

Cell Line↗

Genomic maps and comparative analysis of histone modifications in human and mouse.

We mapped histone H3 lysine 4 di- and trimethylation and lysine 9/14 acetylation across the nonrepetitive portions of human chromosomes 21 and 22 and compared patterns of lysine 4 dimethylation for several orthologous human and mouse loci. Both chromosomes show punctate sites enriched for modified histones. Sites showing trimethylation correlate with transcription starts, while those showing mainly dimethylation occur elsewhere in the vicinity of active genes. Punctate methylation patterns are also evident at the cytokine and IL-4 receptor loci. The Hox clusters present a strikingly different picture, with broad lysine 4-methylated regions that overlay multiple active genes. We suggest these regions represent active chromatin domains required for the maintenance of Hox gene expression. Methylation patterns at orthologous loci are strongly conserved between human and mouse even though many methylated sites do not show sequence conservation notably higher than background. This suggests that the DNA elements that direct the methylation represent only a small fraction of the region or lie at some distance from the site.

Acetylation↗

Examples of the complex architecture of the human transcriptome revealed by RACE and high-density tiling arrays.

Recently, we mapped the sites of transcription across approximately 30% of the human genome and elucidated the structures of several hundred novel transcripts. In this report, we describe a novel combination of techniques including the rapid amplification of cDNA ends (RACE) and tiling array technologies that was used to further characterize transcripts in the human transcriptome. This technical approach allows for several important pieces of information to be gathered about each array-detected transcribed region, including strand of origin, start and termination positions, and the exonic structures of spliced and unspliced coding and noncoding RNAs. In this report, the structures of transcripts from 14 transcribed loci, representing both known genes and unannotated transcripts taken from the several hundred randomly selected unannotated transcripts described in our previous work are represented as examples of the complex organization of the human transcriptome. As a consequence of this complexity, it is not unusual that a single base pair can be part of an intricate network of multiple isoforms of overlapping sense and antisense transcripts, the majority of which are unannotated. Some of these transcripts follow the canonical splicing rules, whereas others combine the exons of different genes or represent other types of noncanonical transcripts. These results have important implications concerning the correlation of genotypes to phenotypes, the regulation of complex interlaced transcriptional patterns, and the definition of a gene.

Cell Line↗

Unbiased mapping of transcription factor binding sites along human chromosomes 21 and 22 points to widespread regulation of noncoding RNAs.

Using high-density oligonucleotide arrays representing essentially all nonrepetitive sequences on human chromosomes 21 and 22, we map the binding sites in vivo for three DNA binding transcription factors, Sp1, cMyc, and p53, in an unbiased manner. This mapping reveals an unexpectedly large number of transcription factor binding site (TFBS) regions, with a minimal estimate of 12,000 for Sp1, 25,000 for cMyc, and 1600 for p53 when extrapolated to the full genome. Only 22% of these TFBS regions are located at the 5' termini of protein-coding genes while 36% lie within or immediately 3' to well-characterized genes and are significantly correlated with noncoding RNAs. A significant number of these noncoding RNAs are regulated in response to retinoic acid, and overlapping pairs of protein-coding and noncoding RNAs are often coregulated. Thus, the human genome contains roughly comparable numbers of protein-coding and noncoding genes that are bound by common transcription factors and regulated by common environmental signals.

Amino Acid Motifs↗

Novel RNAs identified from an in-depth analysis of the transcriptome of human chromosomes 21 and 22.

In this report, we have achieved a richer view of the transcriptome for Chromosomes 21 and 22 by using high-density oligonucleotide arrays on cytosolic poly(A)(+) RNA. Conservatively, only 31.4% of the observed transcribed nucleotides correspond to well-annotated genes, whereas an additional 4.8% and 14.7% correspond to mRNAs and ESTs, respectively. Approximately 85% of the known exons were detected, and up to 21% of known genes have only a single isoform based on exon-skipping alternative expression. Overall, the expression of the well-characterized exons falls predominately into two categories, uniquely or ubiquitously expressed with an identifiable proportion of antisense transcripts. The remaining observed transcription (49.0%) was outside of any known annotation. These novel transcripts appear to be more cell-line-specific and have lower and less variation in expression than the well-characterized genes. Novel transcripts were further characterized based on their distance to annotations, transcript size, coding capacity, and identification as antisense to intronic sequences. By RT-PCR, 126 novel transcripts were independently verified, resulting in a 65% verification rate. These observations strongly support the argument for a re-evaluation of the total number of human genes and an alternative term for "gene" to encompass these growing, novel classes of RNA transcripts in the human genome.

Cell Line↗