PubMed · 10977067
Regulatory element detection using a probabilistic segmentation model.
Abstract
The availability of genome-wide mRNA expression data for organisms whose genome is fully sequenced provides a unique data set from which to decipher how transcription is regulated by the upstream control region of a gene. A new algorithm is presented which decomposes DNA sequence into the most probable "dictionary" of motifs or words. Identification of words is based on a probabilistic segmentation model in which the significance of longer words is deduced from the frequency of shorter words of various length. This eliminates the need for a separate set of reference data to define probabilities, and genome-wide applications are therefore possible. For the 6,000 upstream regulatory regions in the yeast genome, the 500 strongest motifs from a dictionary of size 1,200 match at a significance level of 15 standard deviations to a database of cis-regulatory elements. Analysis of sets of genes such as those up-regulated during sporulation reveals many new putative regulatory sites in addition to identifying previously known sites.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
H J Bussemaker, H Li, E D Siggia. 2000. Regulatory element detection using a probabilistic segmentation model.. https://pubmed.ncbi.nlm.nih.gov/10977067/
Cite the original work for its findings. Save a collection to share your selection of sources.