Search PubMed⌕ Search

PubMed · 16706714

Multiple testing methods for ChIP-Chip high density oligonucleotide array data.

Abstract

Cawley et al. (2004) have recently mapped the locations of binding sites for three transcription factors along human chromosomes 21 and 22 using ChIP-Chip experiments. ChIP-Chip experiments are a new approach to the genomewide identification of transcription factor binding sites and consist of chromatin (Ch) immunoprecipitation (IP) of transcription factor-bound genomic DNA followed by high density oligonucleotide hybridization (Chip) of the IP-enriched DNA. We investigate the ChIP-Chip data structure and propose methods for inferring the location of transcription factor binding sites from these data. The proposed methods involve testing for each probe whether it is part of a bound sequence using a scan statistic that takes into account the spatial structure of the data. Different multiple testing procedures are considered for controlling the familywise error rate and false discovery rate. A nested-Bonferroni adjustment, which is more powerful than the traditional Bonferroni adjustment when the test statistics are dependent, is discussed. Simulation studies show that taking into account the spatial structure of the data substantially improves the sensitivity of the multiple testing procedures. Application of the proposed methods to ChIP-Chip data for transcription factor p53 identified many potential target binding regions along human chromosomes 21 and 22. Among these identified regions, 18% fall within a 3 kb vicinity of the 5'UTR of a known gene or CpG island and 31% fall between the codon start site and the codon end site of a known gene but not inside an exon. More than half of these potential target sequences contain the p53 consensus binding site or very close matches to it. Moreover, these target segments include the 13 experimentally verified p53 binding regions of Cawley et al. (2004), as well as 49 additional regions that show higher hybridization signal than these 13 experimentally verified regions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sündüz Keleş, Mark J van der Laan, Sandrine Dudoit, Simon E Cawley. 2006. Multiple testing methods for ChIP-Chip high density oligonucleotide array data.. https://doi.org/10.1089/cmb.2006.13.579

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries.

Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5'-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5'-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5'-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri. Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism's own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri. The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

5' Untranslated Regions↗

An expanded realm of anti-CRISPR-associated proteins and regulatory mechanisms.

Many bacteriophages encode anti-CRISPR (Acr) proteins that inhibit bacterial CRISPR-Cas immune systems. Rapid acr gene expression upon phage entry enables CRISPR-Cas neutralization but can impact phage fitness if unregulated. Therefore, Acr production is often controlled by distinct families of co-encoded anti-CRISPR-associated (Aca) proteins, which are usually helix-turn-helix (HTH) regulators that bind DNA within acr-aca operon promoters. Previously, we demonstrated that the Aca2 family additionally represses Acr production translationally by binding structured RNA motifs within the 5' untranslated region (UTR) of the acr-aca mRNA. Here, through systematic bioinformatic analyses, we provide evidence of structured RNA motifs in the 5' UTRs of operons encoding members of other Aca families and show that Aca1 also specifically binds its cognate RNA motif. Additionally, many Aca proteins are predicted to regulate not only their own but also adjacent operons with potential anti-defence genes. Indeed, we show that Aca14, newly identified in this study, represses two predicted anti-defence operons. Aca14 is a ribbon-helix-helix domain protein, revealing regulatory diversity beyond the canonical HTH Aca family members. Collectively, our findings expand our understanding of acr regulation in mobile genetic elements and reveal novel mechanisms by which phages fine-tune anti-defence gene expression.

5' Untranslated Regions↗

Functional analysis of stem-loop structures within the SARS-CoV-2 5' untranslated region using a plasmid-based reporter system.

The 5' untranslated region (5'UTR) of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) contains highly conserved stem-loop structures that regulate viral gene expression. This study investigated the functional contributions of selected 5'UTR stem-loop elements to reporter gene expression using a plasmid-based mammalian expression system. Five constructs were tested using a non-integrating plasmid: the wild-type (WT) 5'UTR fused to GFP under the CMV promoter, and four deletion variants (&#x394;B, &#x394;C, &#x394;D, and &#x394;E) corresponding to deletions of stem-loop 4 (SL4), SL4.5, SL5, and SL5a, respectively. Following transfection into HEK293 cells, GFP fluorescence was quantified using a fluorescence microplate reader, and relative GFP transcript abundance was assessed by RT-qPCR. Deletion of SL4 (&#x394;B) resulted in marked reduction in both fluorescence and relative transcript abundance compared to WT construct, indicating substantially reduced reporter gene expression. In contrast, deletion of SL4.5, SL5, or SL5a did not produce the pronounced reduction observed for &#x394;B, although descriptive RT-qPCR analysis indicated differences in relative transcript abundance among these variants. Statistical analysis of fluorescence data demonstrated significant differences among constructs (one-way ANOVA, p&#x2009;<&#x2009;0.05). Because the reporter assay was based on plasmid expression, the observed differences likely reflect combined contributions from transcription, transcript abundance, RNA stability, and translation rather than translation alone. These findings demonstrate that the SL4 region contributes substantially to reporter gene expression in this experimental system, whereas the remaining stem-loop regions examined exert comparatively modest effects. This study provides additional insight into the functional organization of the SARS-CoV-2 5'UTR and establishes a framework for future investigations aimed at distinguished the transcriptional, post-transcriptional, and translational contributions of individual RNA structural elements.

5' Untranslated Regions↗