Search PubMed⌕ Search

Biomedical subjects

Martin C Frith

Publications and source records attributed to Martin C Frith.

10 recordsLinked to original sources

CARRIE web service: automated transcriptional regulatory network inference and interactive analysis.

We present an intuitive and interactive web service for CARRIE (Computational Ascertainment of Regulatory Relationships Inferred from Expression). CARRIE is a computational method that analyzes microarray and promoter sequence data to infer a transcriptional regulatory network from the response to a specific stimulus. This service displays an interactive graph of the inferred network and provides easy access to the evidence for the involvement of each gene in the network. We provide functionality to include network data in KEGG XML (KGML) format in this graph. Our service also provides Gene Ontology annotation to aid the user in forming hypotheses about the role of each gene in the cellular response. The CARRIE web service is freely available at http://zlab.bu.edu/CARRIE-web.

Binding Sites↗

MotifViz: an analysis and visualization tool for motif discovery.

Detecting overrepresented known transcription factor binding motifs in a set of promoter sequences of co-regulated genes has become an important approach to deciphering transcriptional regulatory mechanisms. In this paper, we present an interactive web server, MotifViz, for three motif discovery programs, Clover, Rover and Motifish, covering most available flavors of algorithms for achieving this goal. For comparison, we have also implemented the simple motif-matching program Possum. MotifViz provides uniform and intuitive input and output formats for all four programs. It can be accessed at http://biowulf.bu.edu/MotifViz.

Algorithms↗

Genomic targets of nuclear estrogen receptors.

Estrogen influences the physiology of many target tissues in both women and men. The long-term effects of estrogen are mediated predominantly by nuclear estrogen receptors (ERs) functioning as DNA-binding transcription factors. Tissue-specific responses to estrogen therefore result from regulation of different sets of genes. However, it remains perplexing as to what regulatory sequence contexts specify distinct genomic responses. First, this review classifies estrogen response sequences in mammalian target genes. Of note, around one third of known human target genes associate only indirectly with ER, through intermediary transcription factor(s). Then, computational approaches are presented both for refining direct ER-binding sites and for formulating hypotheses regarding the overall genomic expression pattern. Surprisingly, limited evolutionary conservation of specific estrogen-responsive sites is observed between human and mouse. Finally, consideration of the cellular functions of regulated human genes suggests links between particular biological roles and specific types of estrogen response elements, although with the important caveat that only a restricted set of target genes is available. These analyses support the view that specific, hormone-driven gene expression programs can result from the interplay of environmental and cellular cues with the distinct types of estrogen-response sequences.

Animals↗

Characterization of genomic organization of the adenosine A2A receptor gene by molecular and bioinformatics analyses.

The adenosine A(2A) receptor (A(2A)R) is abundantly expressed in brain and emerging as an important therapeutic target for Parkinson's disease and potentially other neuropsychiatric disorders. To understand the molecular mechanisms of A(2A)R gene expression, we have characterized the genomic organization of the mouse and human A(2A)R genes by molecular and bioinformatic analyses. Three new exons (m1A, m1B and m1C) encoding the 5' untranslated regions (5'-UTRs) of mouse A(2A)R mRNA were identified by rapid amplification of 5' cDNA end (5' RACE), RT-PCR analysis and genome sequence analyses. Similar bioinformatics analysis also suggested six variants of the non-coding "exon 1" (h1A, h1B, h1C, h1D, h1E and h1F) in the human A(2A)R gene, which were confirmed by RT-PCR analysis, while three of the human exon 1 variants (h1D, h1E and h1F) were likewise verified by 5' oligonucleotide capping analysis suggesting multiple transcription start sites. Importantly, RT-PCR and quantitative PCR analysis demonstrated that the A(2A)R transcripts with different exon 1 variants displayed tissue-specific expression patterns. For instance, the mouse exon m1A mRNA was detected only in brain (specifically striatum) and the human exon h1D mRNA in lymphoreticular system. Furthermore, the determination of the three new transcription start sites of human A(2A)R gene by 5' oligonucleotide capping and bioinformatics analyses led to the identification of three corresponding promoter regions which contain several important cis elements, providing additional target for further molecular dissection of A(2A)R gene expression. Finally, our analysis indicates that A(2A)R mRNA and a novel transcript partially overlapping with the 3' exon h3, but in opposite orientation to the A(2A)R gene, could conceivably form duplexes to mutually regulate transcript expression. Thus, combined molecular and bioinformatics analyses revealed a new A(2A)R genomic structure, with conserved coding exons 2 and 3 and divergent, tissue-specific exon 1 variants encoding for 5'-UTR. This raises the possibility of generating multiple tissue-specific A(2A)R mRNA species by alternative promoters with varying regulatory susceptibility.

Animals↗

Detection of functional DNA motifs via statistical over-representation.

The interaction of proteins with DNA recognition motifs regulates a number of fundamental biological processes, including transcription. To understand these processes, we need to know which motifs are present in a sequence and which factors bind to them. We describe a method to screen a set of DNA sequences against a precompiled library of motifs, and assess which, if any, of the motifs are statistically over- or under-represented in the sequences. Over-represented motifs are good candidates for playing a functional role in the sequences, while under-representation hints that if the motif were present, it would have a harmful dysregulatory effect. We apply our method (implemented as a computer program called Clover) to dopamine-responsive promoters, sequences flanking binding sites for the transcription factor LSF, sequences that direct transcription in muscle and liver, and Drosophila segmentation enhancers. In each case Clover successfully detects motifs known to function in the sequences, and intriguing and testable hypotheses are made concerning additional motifs. Clover compares favorably with an ab initio motif discovery algorithm based on sequence alignment, when the motif library includes only a homolog of the factor that actually regulates the sequences. It also demonstrates superior performance over two contingency table based over-representation methods. In conclusion, Clover has the potential to greatly accelerate characterization of signals that regulate transcription.

Animals↗

Site2genome: locating short DNA sequences in whole genomes.

SUMMARY: Many biological papers describe short, functional DNA sites without specifying their exact positions in the genome. We have developed a Web server that automates the tedious task of locating such sites in eukaryotic genomes, thus giving access to the context of rich annotations that are increasingly available for genome sequences. AVAILABILITY: http://zlab.bu.edu/site2genome/

Algorithms↗

Finding functional sequence elements by multiple local alignment.

Algorithms that detect and align locally similar regions of biological sequences have the potential to discover a wide variety of functional motifs. Two theoretical contributions to this classic but unsolved problem are presented here: a method to determine the width of the aligned motif automatically; and a technique for calculating the statistical significance of alignments, i.e. an assessment of whether the alignments are stronger than those that would be expected to occur by chance among random, unrelated sequences. Upon exploring variants of the standard Gibbs sampling technique to optimize the alignment, we discovered that simulated annealing approaches perform more efficiently. Finally, we conduct failure tests by applying the algorithm to increasingly difficult test cases, and analyze the manner of and reasons for eventual failure. Detection of transcription factor-binding motifs is limited by the motifs' intrinsic subtlety rather than by inadequacy of the alignment optimization procedure.

Algorithms↗

Cluster-Buster: Finding dense clusters of motifs in DNA sequences.

The signals that determine activation and repression of specific genes in response to appropriate stimuli are one of the most important, but least understood, types of information encoded in genomic DNA. The nucleotide sequence patterns, or motifs, preferentially bound by various transcription factors have been collected in databases. However, these motifs appear to be individually too short and degenerate to enable detection of functional enhancer and silencer elements within a large genome. Several groups have proposed that dense clusters of motifs may diagnose regulatory regions more accurately. Cluster-Buster is the third incarnation of our software for finding clusters of pre-specified motifs in DNA sequences. We offer a Cluster-Buster web server at http://zlab.bu.edu/cluster-buster/.

Binding Sites↗

A ubiquitous and conserved signal for RNA localization in chordates.

During oogenesis in Xenopus laevis, several RNAs that localize to the vegetal cortex via one of three temporally defined pathways have been identified. Although individual mRNAs utilize only one pathway, there is functional overlap and apparent continuity between them, suggesting that common cis-acting sequences may exist. Because previous work with the Vg1 mRNA revealed that short nontandem repeats are important for localization, we developed a new computer program, called REPFIND, to expedite the identification of repeated motifs in other localized RNAs. Here we show that clusters of short CAC-containing motifs characterize the localization elements (LEs) of virtually all mRNAs localized to the vegetal cortex of Xenopus oocytes. A search for this signal in GenBank [9] resulted in the identification of new localized mRNAs, demonstrating the applicability of REPFIND to predict localized RNAs. CAC-rich LEs are also found in ascidians and other vertebrates, indicating that these cis regulatory elements are conserved in chordates. Interestingly, biochemical evidence shows that distinct CAC-containing motifs have different functions in the localization process. Thus, clusters of CAC-containing motifs are a ubiquitous signal for RNA localization and can signal localization in a variety of pathways through slight variations in sequence composition.

Animals↗

Statistical significance of clusters of motifs represented by position specific scoring matrices in nucleotide sequences.

The human genome encodes the transcriptional control of its genes in clusters of cis-elements that constitute enhancers, silencers and promoter signals. The sequence motifs of individual cis- elements are usually too short and degenerate for confident detection. In most cases, the requirements for organization of cis-elements within these clusters are poorly understood. Therefore, we have developed a general method to detect local concentrations of cis-element motifs, using predetermined matrix representations of the cis-elements, and calculate the statistical significance of these motif clusters. The statistical significance calculation is highly accurate not only for idealized, pseudorandom DNA, but also for real human DNA. We use our method 'cluster of motifs E-value tool' (COMET) to make novel predictions concerning the regulation of genes by transcription factors associated with muscle. COMET performs comparably with two alternative state-of-the-art techniques, which are more complex and lack E-value calculations. Our statistical method enables us to clarify the major bottleneck in the hard problem of detecting cis-regulatory regions, which is that many known enhancers do not contain very significant clusters of the motif types that we search for. Thus, discovery of additional signals that belong to these regulatory regions will be the key to future progress.

Algorithms↗