Search PubMed⌕ Search

Biomedical subjects

Yutao Fu

Publications and source records attributed to Yutao Fu.

7 recordsLinked to original sources

Assessing computational tools for the discovery of transcription factor binding sites.

The prediction of regulatory elements is a problem where computational methods offer great hope. Over the past few years, numerous tools have become available for this task. The purpose of the current assessment is twofold: to provide some guidance to users regarding the accuracy of currently available tools in various settings, and to provide a benchmark of data sets for assessing future tools.

Amino Acid Motifs↗

Improvement of TRANSFAC matrices using multiple local alignment of transcription factor binding site sequences.

This paper describes a novel approach to constructing Position-Specific Weight Matrices (PWMs) based on the transcription factor binding site (TFBS) data provide by the TRANSFAC database and comparison of the newly generated PWMs with the original TRANSFAC matrices. Multiple local sequence alignment was performed on the TFBSs of each transcription factor. Several different alignment programs were tested and their matrices were compared to the original TRANSFAC matrices. One of the alignment programs, GLAM, produced comparable matrices in terms of the average ranking of true positive sites across the whole test set of sequences.

Algorithms↗

MotifViz: an analysis and visualization tool for motif discovery.

Detecting overrepresented known transcription factor binding motifs in a set of promoter sequences of co-regulated genes has become an important approach to deciphering transcriptional regulatory mechanisms. In this paper, we present an interactive web server, MotifViz, for three motif discovery programs, Clover, Rover and Motifish, covering most available flavors of algorithms for achieving this goal. For comparison, we have also implemented the simple motif-matching program Possum. MotifViz provides uniform and intuitive input and output formats for all four programs. It can be accessed at http://biowulf.bu.edu/MotifViz.

Algorithms↗

SeqVISTA: a new module of integrated computational tools for studying transcriptional regulation.

Transcriptional regulation is one of the most basic regulatory mechanisms in the cell. The accumulation of multiple metazoan genome sequences and the advent of high-throughput experimental techniques have motivated the development of a large number of bioinformatics methods for the detection of regulatory motifs. The regulatory process is extremely complex and individual computational algorithms typically have very limited success in genome-scale studies. Here, we argue the importance of integrating multiple computational algorithms and present an infrastructure that integrates eight web services covering key areas of transcriptional regulation. We have adopted the client-side integration technology and built a consistent input and output environment with a versatile visualization tool named SeqVISTA. The infrastructure will allow for easy integration of gene regulation analysis software that is scattered over the Internet. It will also enable bench biologists to perform an arsenal of analysis using cutting-edge methods in a familiar environment and bioinformatics researchers to focus on developing new algorithms without the need to invest substantial effort on complex pre- or post-processors. SeqVISTA is freely available to academic users and can be launched online at http://zlab.bu.edu/SeqVISTA/web.jnlp, provided that Java Web Start has been installed. In addition, a stand-alone version of the program can be downloaded and run locally. It can be obtained at http://zlab.bu.edu/SeqVISTA.

Algorithms↗

Detection of functional DNA motifs via statistical over-representation.

The interaction of proteins with DNA recognition motifs regulates a number of fundamental biological processes, including transcription. To understand these processes, we need to know which motifs are present in a sequence and which factors bind to them. We describe a method to screen a set of DNA sequences against a precompiled library of motifs, and assess which, if any, of the motifs are statistically over- or under-represented in the sequences. Over-represented motifs are good candidates for playing a functional role in the sequences, while under-representation hints that if the motif were present, it would have a harmful dysregulatory effect. We apply our method (implemented as a computer program called Clover) to dopamine-responsive promoters, sequences flanking binding sites for the transcription factor LSF, sequences that direct transcription in muscle and liver, and Drosophila segmentation enhancers. In each case Clover successfully detects motifs known to function in the sequences, and intriguing and testable hypotheses are made concerning additional motifs. Clover compares favorably with an ab initio motif discovery algorithm based on sequence alignment, when the motif library includes only a homolog of the factor that actually regulates the sequences. It also demonstrates superior performance over two contingency table based over-representation methods. In conclusion, Clover has the potential to greatly accelerate characterization of signals that regulate transcription.

Animals↗

Gene expression module discovery using gibbs sampling.

Recent advances in high throughput profiling of gene expression have catalyzed an explosive growth in functional genomics aimed at the elucidation of genes that are differentially expressed in various tissue or cell types across a range of experimental conditions. These studies can lead to the identification of diagnostic genes, classification of genes into functional categories, association of genes with regulatory pathways, and clustering of genes into modules that are potentially co-regulated by a group of transcription factors. Traditional clustering methods such as hierarchical clustering or principal component analysis are difficult to deploy effectively for several of these tasks since genes rarely exhibit similar expression pattern across a wide range of conditions. Bi-clustering of gene expression data is a promising methodology for identification of gene groups that show a coherent expression profile across a subset of conditions. This methodology can be a first step towards the discovery of co-regulated and co-expressed genes or modules. Although bi-clustering (also called block clustering) was introduced in statistics in 1974 few robust and efficient solutions exist for extracting gene expression modules in microarray data. In this paper, we propose a simple but promising new approach for bi-clustering based on a Gibbs sampling paradigm. Our algorithm is implemented in the program GEMS (Gene Expression Module Sampler). GEMS has been tested on synthetic data generated to evaluate the effect of noise on the performance of the algorithm as well as on published leukemia datasets. In our preliminary studies comparing GEMS with other bi-clustering software we show that GEMS is a reliable, flexible and computationally efficient approach for bi-clustering gene expression data.

Cluster Analysis↗

Protistan grazing analysis by flow cytometry using prey labeled by in vivo expression of fluorescent proteins.

Selective grazing by protists can profoundly influence bacterial community structure, and yet direct, quantitative observation of grazing selectivity has been difficult to achieve. In this investigation, flow cytometry was used to study grazing by the marine heterotrophic flagellate Paraphysomonas imperforata on live bacterial cells genetically modified to express the fluorescent protein markers green fluorescent protein (GFP) and red fluorescent protein (RFP). Broad-host-range plasmids were constructed that express fluorescent proteins in three bacterial prey species, Escherichia coli, Enterobacter aerogenes, and Pseudomonas putida. Micromonas pusilla, an alga with red autofluorescence, was also used as prey. Predator-prey interactions were quantified by using a FACScan flow cytometer and analyzed by using a Perl program described here. Grazing preference of P. imperforata was influenced by prey type, size, and condition. In competitive feeding trials, P. imperforata consumed algal prey at significantly lower rates than FP (fluorescent protein)-labeled bacteria of similar or different size. Within-species size selection was also observed, but only for P. putida, the largest prey species examined; smaller cells of P. putida were grazed preferentially. No significant difference in clearance rate was observed between GFP- and RFP-labeled strains of the same prey species or between wild-type and GFP-labeled strains. In contrast, the common chemical staining method, 5-(4,6-dichloro-triazin-2-yl)-amino fluorescein hydrochloride, depressed clearance rates for bacterial prey compared to unlabeled or RFP-labeled cells.

Animals↗