Search PubMed⌕ Search

Biomedical subjects

Gary D Stormo

Publications and source records attributed to Gary D Stormo.

At least 19 recordsLinked to original sources

Using mRNAs lengths to accurately predict the alternatively spliced gene products in Caenorhabditis elegans.

MOTIVATION: Computational gene prediction methods are an important component of whole genome analyses. While ab initio gene finders have demonstrated major improvements in accuracy, the most reliable methods are evidence-based gene predictors. These algorithms can rely on several different sources of evidence including predictions from multiple ab initio gene finders, matches to known proteins, sequence conservation and partial cDNAs to predict the final product. Despite the success of these algorithms, prediction of complete gene structures, especially for alternatively spliced products, remains a difficult task. RESULTS: LOCUS (Length Optimized Characterization of Unknown Spliceforms) is a new evidence-based gene finding algorithm which integrates a length-constraint into a dynamic programming-based framework for prediction of gene products. On a Caenorhabditis elegans test set of alternatively spliced internal exons, its performance exceeds that of current ab initio gene finders and in most cases can accurately predict the correct form of all the alternative products. As the length information used by the algorithm can be obtained in a high-throughput fashion, we propose that integration of such information into a gene-prediction pipeline is feasible and doing so may improve our ability to fully characterize the complete set of mRNAs for a genome. AVAILABILITY: LOCUS is available from http://ural.wustl.edu/software.html

Algorithms↗

An improved map of conserved regulatory sites for Saccharomyces cerevisiae.

BACKGROUND: The regulatory map of a genome consists of the binding sites for proteins that determine the transcription of nearby genes. An initial regulatory map for S. cerevisiae was recently published using six motif discovery programs to analyze genome-wide chromatin immunoprecipitation data for 203 transcription factors. The programs were used to identify sequence motifs that were likely to correspond to the DNA-binding specificity of the immunoprecipitated proteins. We report improved versions of two conservation-based motif discovery algorithms, PhyloCon and Converge. Using these programs, we create a refined regulatory map for S. cerevisiae by reanalyzing the same chromatin immunoprecipitation data. RESULTS: Applying the same conservative criteria that were applied in the original study, we find that PhyloCon and Converge each separately discover more known specificities than the combination of all six programs in the previous study. Combining the results of PhyloCon and Converge, we discover significant sequence motifs for 36 transcription factors that were previously missed. The new set of motifs identifies 636 more regulatory interactions than the previous one. The new network contains 28% more regulatory interactions among transcription factors, evidence of greater cross-talk between regulators. CONCLUSION: Combining two complementary computational strategies for conservation-based motif discovery improves the ability to identify the specificity of transcriptional regulators from genome-wide chromatin immunoprecipitation data. The increased sensitivity of these methods significantly expands the map of yeast regulatory sites without the need to alter any of the thresholds for statistical significance. The new map of regulatory sites reveals a more elaborate and complex view of the yeast genetic regulatory network than was observed previously.

Algorithms↗

A systematic model to predict transcriptional regulatory mechanisms based on overrepresentation of transcription factor binding profiles.

An important aspect of understanding a biological pathway is to delineate the transcriptional regulatory mechanisms of the genes involved. Two important tasks are often encountered when studying transcription regulation, i.e., (1) the identification of common transcriptional regulators of a set of coexpressed genes; (2) the identification of genes that are regulated by one or several transcription factors. In this study, a systematic and statistical approach was taken to accomplish these tasks by establishing an integrated model considering all of the promoters and characterized transcription factors (TFs) in the genome. A promoter analysis pipeline (PAP) was developed to implement this approach. PAP was tested using coregulated gene clusters collected from the literature. In most test cases, PAP identified the transcription regulators of the input genes accurately. When compared with chromatin immunoprecipitation experiment data, PAP's predictions are consistent with the experimental observations. When PAP was used to analyze one published expression-profiling data set and two novel coregulated gene sets, PAP was able to generate biologically meaningful hypotheses. Therefore, by taking a systematic approach of considering all promoters and characterized TFs in our model, we were able to make more reliable predictions about the regulation of gene expression in mammalian organisms.

Animals↗

Target selectivity of vertebrate notch proteins. Collaboration between discrete domains and CSL-binding site architecture determines activation probability.

All four mammalian Notch proteins interact with a single DNA-binding protein (RBP-jkappa), yet they are not equivalent in activating target genes. Parallel assays of three Notch-responsive promoters in several cell lines revealed that relative activation strength is dependent on protein module and promoter context more than the cellular context. Each Notch protein reads binding site orientation and distribution on the promoter differently; Notch1 performs extremely well on paired sites, and Notch3 prefers single sites in conjunction with a proximal zinc finger transcription factor. Although head-head sites can elicit a Notch response on their own, use of CBS (CSL binding site) in tail-tail orientation is context-dependent. Bias for specific DNA elements is achieved by interplay between the N-terminal RAM (RBP-jkappa-associated molecule/ankyrin region), which interprets CBS proximity and orientation, and the C-terminal transactivation domain that interacts specifically with the transcription machinery or nearby factors. To confirm the prediction that modular design underscores the evolution of functional divergence between Notch proteins, we generated a synthetic Notch protein (Notch1 ankyrin with Notch3 transactivation domain) that displayed superior signaling strength on the hes5 promoter. Consistent with the prediction that "preferred" targets (Hes1) should respond faster and at lower Notch concentration than other targets, we showed that Hes5-GFP was extinguished fast and recovered slowly, whereas Hes1-GFP was inhibited late and recovered quickly after a pulse of DAPT in metanephroi cultures.

Animals↗

Identifying the conserved network of cis-regulatory sites of a eukaryotic genome.

A major focus of genome research has been to decipher the cis-regulatory code that governs complex transcriptional regulation. We report a computational approach for identifying conserved regulatory motifs of an organism directly from whole genome sequences of several related species without reliance on additional information. We first construct phylogenetic profiles for each promoter, then use a BLAST-like algorithm to efficiently search through the entire profile space of all of the promoters in the genome to identify conserved motifs and the promoters that contain them. Statistical significance is estimated by modified Karlin-Altschul statistics. We applied this approach to the analysis of 3,524 Saccharomyces cerevisiae promoters and identified a highly organized regulatory network involving 3,315 promoters and 296 motifs. This network includes nearly all of the currently known motifs and covers >90% of known transcription factor binding sites. Most of the predicted coregulated gene clusters in the network have additional supporting evidence. Theoretical analysis suggests that our algorithm should be applicable to much larger genomes, such as the human genome, without reaching its statistical limitation.

Algorithms↗

Direct, androgen receptor-mediated regulation of the FKBP5 gene via a distal enhancer element.

Androgen signaling via the androgen receptor (AR) transcription factor is crucial to normal prostate homeostasis and prostate tumorigenesis. Current models of AR function are predominantly based on studies of prostate-specific antigen regulation in androgen-responsive cell lines. To expand on these in vitro paradigms, we used the mouse prostate to elucidate the mechanisms through which AR regulates another direct target, FKBP5, in vivo. FKBP5 encodes an immunophilin that has been previously implicated in glucocorticoid and progestin signaling pathways and that likely influences prostate physiology in the presence of androgens. In this work, we show that androgens directly regulate FKBP5 via an interaction between the AR and a distal enhancer located 65 kb downstream of the transcription start site in the fifth intron of the FKBP5 gene. We have found that AR selectively recruits cAMP response element-binding protein to this enhancer. These interactions, in turn, result in chromatin remodeling that affects the enhancer proper but not the FKBP5 locus as a whole. Furthermore, in contrast to prostate-specific antigen-regulatory mechanisms, we show that transactivation of the FKBP5 gene does not rely on a single looping complex to mediate communication between the distal enhancer and proximal promoter. Rather, the distal enhancer complex and basal transcription apparatus communicate indirectly with one another, implicating a regulatory mechanism that has not been previously appreciated for AR target genes.

Animals↗

Combining SELEX with quantitative assays to rapidly obtain accurate models of protein-DNA interactions.

Models for the specificity of DNA-binding transcription factors are often based on small amounts of qualitative data and therefore have limited accuracy. In this study we demonstrate a simple and efficient method of affinity chromatography-SELEX followed by a quantitative binding (QuMFRA) assay to rapidly collect the data necessary for more accurate models. Using the zinc finger protein EGR as an e.g. we show that many bindings sites can be obtained efficiently with affinity chromatography-SELEX, but those sequences alone provide a weight matrix model with limited accuracy. Using a QuMFRA assay to determine the quantitative relative affinity for only a subset of the sequences obtained by SELEX leads to a much more accurate model. Application of this method to variants of a transcription factor would allow us to generate a large collection of quantitative data for modeling protein-DNA interactions that could facilitate the determination of recognition codes for different transcription factor families.

Base Sequence↗

Quantitative analysis of EGR proteins binding to DNA: assessing additivity in both the binding site and the protein.

BACKGROUND: Recognition codes for protein-DNA interactions typically assume that the interacting positions contribute additively to the binding energy. While this is known to not be precisely true, an additive model over the DNA positions can be a good approximation, at least for some proteins. Much less information is available about whether the protein positions contribute additively to the interaction. RESULTS: Using EGR zinc finger proteins, we measure the binding affinity of six different variants of the protein to each of six different variants of the consensus binding site. Both the protein and binding site variants include single and double mutations that allow us to assess how well additive models can account for the data. For each protein and DNA alone we find that additive models are good approximations, but over the combined set of data there are context effects that limit their accuracy. However, a small modification to the purely additive model, with only three additional parameters, improves the fit significantly. CONCLUSION: The additive model holds very well for every DNA site and every protein included in this study, but clear context dependence in the interactions was detected. A simple modification to the independent model provides a better fit to the complete data.

Algorithms↗

enoLOGOS: a versatile web tool for energy normalized sequence logos.

enoLOGOS is a web-based tool that generates sequence logos from various input sources. Sequence logos have become a popular way to graphically represent DNA and amino acid sequence patterns from a set of aligned sequences. Each position of the alignment is represented by a column of stacked symbols with its total height reflecting the information content in this position. Currently, the available web servers are able to create logo images from a set of aligned sequences, but none of them generates weighted sequence logos directly from energy measurements or other sources. With the advent of high-throughput technologies for estimating the contact energy of different DNA sequences, tools that can create logos directly from binding affinity data are useful to researchers. enoLOGOS generates sequence logos from a variety of input data, including energy measurements, probability matrices, alignment matrices, count matrices and aligned sequences. Furthermore, enoLOGOS can represent the mutual information of different positions of the consensus sequence, a unique feature of this tool. Another web interface for our software, C2H2-enoLOGOS, generates logos for the DNA-binding preferences of the C2H2 zinc-finger transcription factor family members. enoLOGOS and C2H2-enoLOGOS are accessible over the web at http://biodev.hgen.pitt.edu/enologos/.

Amino Acids↗

Computational technique for improvement of the position-weight matrices for the DNA/protein binding sites.

Position-weight matrices (PWMs) are broadly used to locate transcription factor binding sites in DNA sequences. The majority of existing PWMs provide a low level of both sensitivity and specificity. We present a new computational algorithm, a modification of the Staden-Bucher approach, that improves the PWM. We applied the proposed technique on the PWM of the GC-box, binding site for Sp1. The comparison of old and new PWMs shows that the latter increase both sensitivity and specificity. The statistical parameters of GC-box distribution in promoter regions and in the human genome, as well as in each chromosome, are presented. The majority of commonly used PWMs are the 4-row mononucleotide matrices, although 16-row dinucleotide matrices are known to be more informative. The algorithm efficiently determines the 16-row matrices and preliminary results show that such matrices provide better results than 4-row matrices.

Algorithms↗

Pairwise local structural alignment of RNA sequences with sequence similarity less than 40%.

MOTIVATION: Searching for non-coding RNA (ncRNA) genes and structural RNA elements (eleRNA) are major challenges in gene finding today as these often are conserved in structure rather than in sequence. Even though the number of available methods is growing, it is still of interest to pairwise detect two genes with low sequence similarity, where the genes are part of a larger genomic region. RESULTS: Here we present such an approach for pairwise local alignment which is based on foldalign and the Sankoff algorithm for simultaneous structural alignment of multiple sequences. We include the ability to conduct mutual scans of two sequences of arbitrary length while searching for common local structural motifs of some maximum length. This drastically reduces the complexity of the algorithm. The scoring scheme includes structural parameters corresponding to those available for free energy as well as for substitution matrices similar to RIBOSUM. The new foldalign implementation is tested on a dataset where the ncRNAs and eleRNAs have sequence similarity <40% and where the ncRNAs and eleRNAs are energetically indistinguishable from the surrounding genomic sequence context. The method is tested in two ways: (1) its ability to find the common structure between the genes only and (2) its ability to locate ncRNAs and eleRNAs in a genomic context. In case (1), it makes sense to compare with methods like Dynalign, and the performances are very similar, but foldalign is substantially faster. The structure prediction performance for a family is typically around 0.7 using Matthews correlation coefficient. In case (2), the algorithm is successful at locating RNA families with an average sensitivity of 0.8 and a positive predictive value of 0.9 using a BLAST-like hit selection scheme. AVAILABILITY: The program is available online at http://foldalign.kvl.dk/

Algorithms↗

Making connections between novel transcription factors and their DNA motifs.

The key components of a transcriptional regulatory network are the connections between trans-acting transcription factors and cis-acting DNA-binding sites. In spite of several decades of intense research, only a fraction of the estimated approximately 300 transcription factors in Escherichia coli have been linked to some of their binding sites in the genome. In this paper, we present a computational method to connect novel transcription factors and DNA motifs in E. coli. Our method uses three types of mutually independent information, two of which are gleaned by comparative analysis of multiple genomes and the third one derived from similarities of transcription-factor-DNA-binding-site interactions. The different types of information are combined to calculate the probability of a given transcription-factor-DNA-motif pair being a true pair. Tested on a study set of transcription factors and their DNA motifs, our method has a prediction accuracy of 59% for the top predictions and 85% for the top three predictions. When applied to 99 novel transcription factors and 70 novel DNA motifs, our method predicted 64 transcription-factor-DNA-motif pairs. Supporting evidence for some of the predicted pairs is presented. Functional annotations are made for 23 novel transcription factors based on the predicted transcription-factor-DNA-motif connections.

Algorithms↗

PolyMAPr: programs for polymorphism database mining, annotation, and functional analysis.

Pharmacogenomic and disease-association studies rely on identifying a comprehensive set of polymorphisms within candidate genes. Public SNP databases are a rich source of polymorphism data, but mining them effectively requires overcoming at least four challenges: ensuring accurate annotations for genes and polymorphisms, eliminating both inter- and intra-database redundancy, integrating data from multiple public sources with data generated locally, and prioritizing the variants for further study. PolyMAPr (Polymorphism Mining and Annotation Programs)' was developed to overcome these challenges and to improve the efficiency of database mining and polymorphism annotation. PolyMAPr takes as input a file containing a list of genes to be processed and files containing each annotated gene sequence. Polymorphic sequences obtained from public databases (dbSNP, CGAP, and JSNP) or through local SNP discovery efforts, as well as oligonucleotide sequences (e.g., PCR primers), are mapped to the annotated gene sequences and named according to suggested nomenclature guidelines. The functional effects of nonsynonymous coding-region SNPs (cSNPs) and any variants that might alter exon splicing enhancer (ESE) sites, putative transcription factor binding sites, or intron-exon splice sites are predicted. The output files are accessible though a browser interface. In addition, the results are also provided in Extensible Markup Language (XML) format to facilitate uploading them into a local relational database. PolyMAPr increases the efficiency of mining public databases for genetic variants within candidate genes and provides a mechanism by which data from multiple sources (both public and private) can be uniformly integrated, thereby significantly reducing the effort required to obtain a comprehensive set of polymorphisms for pharmacogenomic and disease-association studies. PolyMAPr can be obtained from http://pharmacogenomics.wustl.edu.

Databases, Nucleic Acid↗

Editing efficiency of a Drosophila gene correlates with a distant splice site selection.

RNA editing and alternative splicing are two processes that increase protein diversity. The relationship between the two processes is not well understood. There are a few examples of correlations between editing and alternative splicing, but these are all nearby effects. A search for alternative splicing among 16 edited genes in Drosophila reveals two novel instances of alternative splicing. In one example where alternative splicing occurs downstream of editing, a strong correlation between editing efficiency and splice site selection is observed. In contrast, when editing occurs downstream of alternative splicing, no correlation is seen. These results suggest some models for the coupling of editing and splicing processes.

Alternative Splicing↗

Procom: a web-based tool to compare multiple eukaryotic proteomes.

UNLABELLED: Each organism has traits that are shared with some, but not all, organisms. Identification of genes needed for a particular trait can be accomplished by a comparative genomics approach using three or more organisms. Genes that occur in organisms without the trait are removed from the set of genes in common among organisms with the trait. To facilitate these comparisons, a web-based server, Procom, was developed to identify the subset of genes that may be needed for a trait. AVAILABILITY: The Procom program is freely available with documentation and examples at http://ural.wustl.edu/~billy/Procom/ CONTACT: billy@ural.wustl.edu.

Animals↗

The neuropeptide pigment-dispersing factor coordinates pacemaker interactions in the Drosophila circadian system.

In Drosophila, the neuropeptide pigment-dispersing factor (PDF) is required to maintain behavioral rhythms under constant conditions. To understand how PDF exerts its influence, we performed time-series immunostainings for the PERIOD protein in normal and pdf mutant flies over 9 d of constant conditions. Without pdf, pacemaker neurons that normally express PDF maintained two markers of rhythms: that of PERIOD nuclear translocation and its protein staining intensity. As a group, however, they displayed a gradual dispersion in their phasing of nuclear translocation. A separate group of non-PDF circadian pacemakers also maintained PERIOD nuclear translocation rhythms without pdf but exhibited altered phase and amplitude of PERIOD staining intensity. Therefore, pdf is not required to maintain circadian protein oscillations under constant conditions; however, it is required to coordinate the phase and amplitude of such rhythms among the diverse pacemakers. These observations begin to outline the hierarchy of circadian pacemaker circuitry in the Drosophila brain.

Animals↗

Quantitative modeling of DNA-protein interactions: effects of amino acid substitutions on binding specificity of the Mnt repressor.

Understanding DNA-protein recognition quantitatively is essential to developing computational algorithms for accurate transcriptional binding site prediction. Using a quantitative, multiple fluorescence, relative affinity (QuMFRA) assay, we determine the binding specificity of 11 different position 6 variants of the Mnt repressor for operators containing all 16 possible dinucleotides at operator positions 16 and 17. We show that the wild-type and all variant proteins interact with the two positions in a non-independent manner, but that a simple independent model provides a close approximation to the true binding affinities. The wild-type His at amino acid 6 is the only protein to prefer the AC sequence of the wild-type operator, whereas most of the variant proteins prefer TA. H6R is unique in having a strong preference for C at position 16. A comparison of the quantitative binding data for all of the protein variants with a model for recognition of the early growth response (EGR) zinc finger family suggests that interactions of Mnt with positions 16 and 17 are similar to interactions of EGR with positions 1 and 2, respectively. This information leads to an augmented model for the interaction of Mnt with its operator.

Amino Acid Substitution↗

ILM: a web server for predicting RNA secondary structures with pseudoknots.

The ILM web server provides a web interface to two algorithms, iterated loop matching and maximum weighted matching, for efficiently predicting RNA secondary structures with pseudoknots. The algorithms can utilize either thermodynamic or comparative information or both, and thus can work on both aligned and individual sequences. Predicted secondary structures are presented in several formats compatible with a variety of existing visualization tools. The service can be accessed at http://cic.cs.wustl.edu/RNA/.

Algorithms↗