Search PubMed⌕ Search

Biomedical subjects

David K Gifford

Publications and source records attributed to David K Gifford.

At least 19 recordsLinked to original sources

Semi-supervised analysis of gene expression profiles for lineage-specific development in the Caenorhabditis elegans embryo.

MOTIVATION: Gene expression profiling is a powerful approach to identify genes that may be involved in a specific biological process on a global scale. For example, gene expression profiling of mutant animals that lack or contain an excess of certain cell types is a common way to identify genes that are important for the development and maintenance of given cell types. However, it is difficult for traditional computational methods, including unsupervised and supervised learning methods, to detect relevant genes from a large collection of expression profiles with high sensitivity and specificity. Unsupervised methods group similar gene expressions together while ignoring important prior biological knowledge. Supervised methods utilize training data from prior biological knowledge to classify gene expression. However, for many biological problems, little prior knowledge is available, which limits the prediction performance of most supervised methods. RESULTS: We present a Bayesian semi-supervised learning method, called BGEN, that improves upon supervised and unsupervised methods by both capturing relevant expression profiles and using prior biological knowledge from literature and experimental validation. Unlike currently available semi-supervised learning methods, this new method trains a kernel classifier based on labeled and unlabeled gene expression examples. The semi-supervised trained classifier can then be used to efficiently classify the remaining genes in the dataset. Moreover, we model the confidence of microarray probes and probabilistically combine multiple probe predictions into gene predictions. We apply BGEN to identify genes involved in the development of a specific cell lineage in the C. elegans embryo, and to further identify the tissues in which these genes are enriched. Compared to K-means clustering and SVM classification, BGEN achieves higher sensitivity and specificity. We confirm certain predictions by biological experiments. AVAILABILITY: The results are available at http://www.csail.mit.edu/~alanqi/projects/BGEN.html.

Algorithms↗

Core transcriptional regulatory circuitry in human hepatocytes.

We mapped the transcriptional regulatory circuitry for six master regulators in human hepatocytes using chromatin immunoprecipitation and high-resolution promoter microarrays. The results show that these regulators form a highly interconnected core circuitry, and reveal the local regulatory network motifs created by regulator-gene interactions. Autoregulation was a prominent theme among these regulators. We found that hepatocyte master regulators tend to bind promoter regions combinatorially and that the number of transcription factors bound to a promoter corresponds with observed gene expression. Our studies reveal portions of the core circuitry of human hepatocytes.

Gene Expression Regulation↗

Control of developmental regulators by Polycomb in human embryonic stem cells.

Polycomb group proteins are essential for early development in metazoans, but their contributions to human development are not well understood. We have mapped the Polycomb Repressive Complex 2 (PRC2) subunit SUZ12 across the entire nonrepeat portion of the genome in human embryonic stem (ES) cells. We found that SUZ12 is distributed across large portions of over two hundred genes encoding key developmental regulators. These genes are occupied by nucleosomes trimethylated at histone H3K27, are transcriptionally repressed, and contain some of the most highly conserved noncoding elements in the genome. We found that PRC2 target genes are preferentially activated during ES cell differentiation and that the ES cell regulators OCT4, SOX2, and NANOG cooccupy a significant subset of these genes. These results indicate that PRC2 occupies a special set of developmental genes in ES cells that must be repressed to maintain pluripotency and that are poised for activation during ES cell differentiation.

Animals↗

Polycomb complexes repress developmental regulators in murine embryonic stem cells.

The mechanisms by which embryonic stem (ES) cells self-renew while maintaining the ability to differentiate into virtually all adult cell types are not well understood. Polycomb group (PcG) proteins are transcriptional repressors that help to maintain cellular identity during metazoan development by epigenetic modification of chromatin structure. PcG proteins have essential roles in early embryonic development and have been implicated in ES cell pluripotency, but few of their target genes are known in mammals. Here we show that PcG proteins directly repress a large cohort of developmental regulators in murine ES cells, the expression of which would otherwise promote differentiation. Using genome-wide location analysis in murine ES cells, we found that the Polycomb repressive complexes PRC1 and PRC2 co-occupied 512 genes, many of which encode transcription factors with important roles in development. All of the co-occupied genes contained modified nucleosomes (trimethylated Lys 27 on histone H3). Consistent with a causal role in gene silencing in ES cells, PcG target genes were de-repressed in cells deficient for the PRC2 component Eed, and were preferentially activated on induction of differentiation. Our results indicate that dynamic repression of developmental pathways by Polycomb complexes may be required for maintaining ES cell pluripotency and plasticity during embryonic development.

Animals↗

Coordinated binding of NF-kappaB family members in the response of human cells to lipopolysaccharide.

The NF-kappaB family of transcription factors plays a critical role in numerous cellular processes, particularly the immune response. Our understanding of how the different NF-kappaB subunits act coordinately to regulate gene expression is based on a limited set of genes. We used genome-scale location analysis to identify targets of all five NF-kappaB proteins before and after stimulation of monocytic cells with bacterial lipopolysaccharide (LPS). In unstimulated cells, p50 and p52 bound to a large number of gene promoters that were also occupied by RNA polymerase II. After LPS stimulation, additional NF-kappaB subunits bound to these genes and to other genes. Genes that became bound by multiple NF-kappaB subunits were the most likely to show increases in RNA polymerase II occupancy and gene expression. This study identifies NF-kappaB target genes, reveals how the different NF-kappaB proteins coordinate their activity, and provides an initial map of the transcriptional regulatory network that underlies the host response to infection.

Genome, Human↗

An improved map of conserved regulatory sites for Saccharomyces cerevisiae.

BACKGROUND: The regulatory map of a genome consists of the binding sites for proteins that determine the transcription of nearby genes. An initial regulatory map for S. cerevisiae was recently published using six motif discovery programs to analyze genome-wide chromatin immunoprecipitation data for 203 transcription factors. The programs were used to identify sequence motifs that were likely to correspond to the DNA-binding specificity of the immunoprecipitated proteins. We report improved versions of two conservation-based motif discovery algorithms, PhyloCon and Converge. Using these programs, we create a refined regulatory map for S. cerevisiae by reanalyzing the same chromatin immunoprecipitation data. RESULTS: Applying the same conservative criteria that were applied in the original study, we find that PhyloCon and Converge each separately discover more known specificities than the combination of all six programs in the previous study. Combining the results of PhyloCon and Converge, we discover significant sequence motifs for 36 transcription factors that were previously missed. The new set of motifs identifies 636 more regulatory interactions than the previous one. The new network contains 28% more regulatory interactions among transcription factors, evidence of greater cross-talk between regulators. CONCLUSION: Combining two complementary computational strategies for conservation-based motif discovery improves the ability to identify the specificity of transcriptional regulators from genome-wide chromatin immunoprecipitation data. The increased sensitivity of these methods significantly expands the map of yeast regulatory sites without the need to alter any of the thresholds for statistical significance. The new map of regulatory sites reveals a more elaborate and complex view of the yeast genetic regulatory network than was observed previously.

Algorithms↗

High-resolution computational models of genome binding events.

Direct physical information that describes where transcription factors, nucleosomes, modified histones, RNA polymerase II and other key proteins interact with the genome provides an invaluable mechanistic foundation for understanding complex programs of gene regulation. We present a method, joint binding deconvolution (JBD), which uses additional easily obtainable experimental data about chromatin immunoprecipitation (ChIP) to improve the spatial resolution of the transcription factor binding locations inferred from ChIP followed by DNA microarray hybridization (ChIP-Chip) data. Based on this probabilistic model of binding data, we further pursue improved spatial resolution by using sequence information. We produce positional priors that link ChIP-Chip data to sequence data by guiding motif discovery to inferred protein-DNA binding sites. We present results on the yeast transcription factors Gcn4 and Mig2 to demonstrate JBD's spatial resolution capabilities and show that positional priors allow computational discovery of the Mig2 motif when a standard approach fails.

Base Sequence↗

A hypothesis-based approach for identifying the binding specificity of regulatory proteins from chromatin immunoprecipitation data.

MOTIVATION: Genome-wide chromatin-immunoprecipitation (ChIP-chip) detects binding of transcriptional regulators to DNA in vivo at low resolution. Motif discovery algorithms can be used to discover sequence patterns in the bound regions that may be recognized by the immunoprecipitated protein. However, the discovered motifs often do not agree with the binding specificity of the protein, when it is known. RESULTS: We present a powerful approach to analyzing ChIP-chip data, called THEME, that tests hypotheses concerning the sequence specificity of a protein. Hypotheses are refined using constrained local optimization. Cross-validation provides a principled standard for selecting the optimal weighting of the hypothesis and the ChIP-chip data and for choosing the best refined hypothesis. We demonstrate how to derive hypotheses for proteins from 36 domain families. Using THEME together with these hypotheses, we analyze ChIP-chip datasets for 14 human and mouse proteins. In all the cases the identified motifs are consistent with the published data with regard to the binding specificity of the proteins.

Algorithms↗

Core transcriptional regulatory circuitry in human embryonic stem cells.

The transcription factors OCT4, SOX2, and NANOG have essential roles in early development and are required for the propagation of undifferentiated embryonic stem (ES) cells in culture. To gain insights into transcriptional regulation of human ES cells, we have identified OCT4, SOX2, and NANOG target genes using genome-scale location analysis. We found, surprisingly, that OCT4, SOX2, and NANOG co-occupy a substantial portion of their target genes. These target genes frequently encode transcription factors, many of which are developmentally important homeodomain proteins. Our data also indicate that OCT4, SOX2, and NANOG collaborate to form regulatory circuitry consisting of autoregulatory and feedforward loops. These results provide new insights into the transcriptional regulation of stem cells and reveal how OCT4, SOX2, and NANOG contribute to pluripotency and self-renewal.

Animals↗

Genome-wide map of nucleosome acetylation and methylation in yeast.

Eukaryotic genomes are packaged into nucleosomes whose position and chemical modification state can profoundly influence regulation of gene expression. We profiled nucleosome modifications across the yeast genome using chromatin immunoprecipitation coupled with DNA microarrays to produce high-resolution genome-wide maps of histone acetylation and methylation. These maps take into account changes in nucleosome occupancy at actively transcribed genes and, in doing so, revise previous assessments of the modifications associated with gene expression. Both acetylation and methylation of histones are associated with transcriptional activity, but the former occurs predominantly at the beginning of genes, whereas the latter can occur throughout transcribed regions. Most notably, specific methylation events are associated with the beginning, middle, and end of actively transcribed genes. These maps provide the foundation for further understanding the roles of chromatin in gene expression and genome maintenance.

Acetylation↗

Global position and recruitment of HATs and HDACs in the yeast genome.

Chromatin regulators play fundamental roles in the regulation of gene expression and chromosome maintenance, but the regions of the genome where most of these regulators function has not been established. We explored the genome-wide occupancy of four different chromatin regulators encoded in Saccharomyces cerevisiae. The results reveal that the histone acetyltransferases Gcn5 and Esa1 are both generally recruited to the promoters of active protein-coding genes. In contrast, the histone deacetylases Hst1 and Rpd3 are recruited to specific sets of genes associated with distinct cellular functions. Our results provide new insights into the association of histone acetyltransferases and histone deacetylases with the yeast genome, and together with previous studies, suggest how these chromatin regulators are recruited to specific regions of the genome.

Acetyltransferases↗

Transcriptional regulatory code of a eukaryotic genome.

DNA-binding transcriptional regulators interpret the genome's regulatory code by binding to specific sequences to induce or repress gene expression. Comparative genomics has recently been used to identify potential cis-regulatory sequences within the yeast genome on the basis of phylogenetic conservation, but this information alone does not reveal if or when transcriptional regulators occupy these binding sites. We have constructed an initial map of yeast's transcriptional regulatory code by identifying the sequence elements that are bound by regulators under various conditions and that are conserved among Saccharomyces species. The organization of regulatory elements in promoters and the environment-dependent use of these elements by regulators are discussed. We find that environment-specific use of regulatory elements predicts mechanistic models for the function of a large population of yeast's transcriptional regulators.

Base Sequence↗

Deconvolving cell cycle expression data with complementary information.

MOTIVATION: In the study of many systems, cells are first synchronized so that a large population of cells exhibit similar behavior. While synchronization can usually be achieved for a short duration, after a while cells begin to lose their synchronization. Synchronization loss is a continuous process and so the observed value in a population of cells for a gene at time t is actually a convolution of its values in an interval around t. Deconvolving the observed values from a mixed population will allow us to obtain better models for these systems and to accurately detect the genes that participate in these systems. RESULTS: We present an algorithm which combines budding index and gene expression data to deconvolve expression profiles. Using the budding index data we first fit a synchronization loss model for the cell cycle system. Our deconvolution algorithm uses this loss model and can also use information from co-expressed genes, making it more robust against noise and missing values. Using expression and budding data for yeast we show that our algorithm is able to reconstruct a more accurate representation when compared with the observed values. In addition, using the deconvolved profiles we are able to correctly identify 15% more cycling genes when compared to a set identified using the observed values. AVAILABILITY: Matlab implementation can be downloaded from the supporting website http://www.cs.cmu.edu/~zivbj/decon/decon.html

Biological Clocks↗

Control of pancreas and liver gene expression by HNF transcription factors.

The transcriptional regulatory networks that specify and maintain human tissue diversity are largely uncharted. To gain insight into this circuitry, we used chromatin immunoprecipitation combined with promoter microarrays to identify systematically the genes occupied by the transcriptional regulators HNF1alpha, HNF4alpha, and HNF6, together with RNA polymerase II, in human liver and pancreatic islets. We identified tissue-specific regulatory circuits formed by HNF1alpha, HNF4alpha, and HNF6 with other transcription factors, revealing how these factors function as master regulators of hepatocyte and islet transcription. Our results suggest how misregulation of HNF4alpha can contribute to type 2 diabetes.

Basic Helix-Loop-Helix Leucine Zipper Transcriptio↗

Computational discovery of gene modules and regulatory networks.

We describe an algorithm for discovering regulatory networks of gene modules, GRAM (Genetic Regulatory Modules), that combines information from genome-wide location and expression data sets. A gene module is defined as a set of coexpressed genes to which the same set of transcription factors binds. Unlike previous approaches that relied primarily on functional information from expression data, the GRAM algorithm explicitly links genes to the factors that regulate them by incorporating DNA binding data, which provide direct physical evidence of regulatory interactions. We use the GRAM algorithm to describe a genome-wide regulatory network in Saccharomyces cerevisiae using binding information for 106 transcription factors profiled in rich medium conditions data from over 500 expression experiments. We also present a genome-wide location analysis data set for regulators in yeast cells treated with rapamycin, and use the GRAM algorithm to provide biological insights into this regulatory network

Algorithms↗

Comparing the continuous representation of time-series expression profiles to identify differentially expressed genes.

We present a general algorithm to detect genes differentially expressed between two nonhomogeneous time-series data sets. As increasing amounts of high-throughput biological data become available, a major challenge in genomic and computational biology is to develop methods for comparing data from different experimental sources. Time-series whole-genome expression data are a particularly valuable source of information because they can describe an unfolding biological process such as the cell cycle or immune response. However, comparisons of time-series expression data sets are hindered by biological and experimental inconsistencies such as differences in sampling rate, variations in the timing of biological processes, and the lack of repeats. Our algorithm overcomes these difficulties by using a continuous representation for time-series data and combining a noise model for individual samples with a global difference measure. We introduce a corresponding statistical method for computing the significance of this differential expression measure. We used our algorithm to compare cell-cycle-dependent gene expression in wild-type and knockout yeast strains. Our algorithm identified a set of 56 differentially expressed genes, and these results were validated by using independent protein-DNA-binding data. Unlike previous methods, our algorithm was also able to identify 22 non-cell-cycle-regulated genes as differentially expressed. This set of genes is significantly correlated in a set of independent expression experiments, suggesting additional roles for the transcription factors Fkh1 and Fkh2 in controlling cellular activity in yeast.

Algorithms↗

K-ary clustering with optimal leaf ordering for gene expression data.

MOTIVATION: A major challenge in gene expression analysis is effective data organization and visualization. One of the most popular tools for this task is hierarchical clustering. Hierarchical clustering allows a user to view relationships in scales ranging from single genes to large sets of genes, while at the same time providing a global view of the expression data. However, hierarchical clustering is very sensitive to noise, it usually lacks of a method to actually identify distinct clusters, and produces a large number of possible leaf orderings of the hierarchical clustering tree. In this paper we propose a new hierarchical clustering algorithm which reduces susceptibility to noise, permits up to k siblings to be directly related, and provides a single optimal order for the resulting tree. RESULTS: We present an algorithm that efficiently constructs a k-ary tree, where each node can have up to k children, and then optimally orders the leaves of that tree. By combining k clusters at each step our algorithm becomes more robust against noise and missing values. By optimally ordering the leaves of the resulting tree we maintain the pairwise relationships that appear in the original method, without sacrificing the robustness. Our k-ary construction algorithm runs in O(n(3)) regardless of k and our ordering algorithm runs in O(4(k)n(3)). We present several examples that show that our k-ary clustering algorithm achieves results that are superior to the binary tree results in both global presentation and cluster identification. AVAILABILITY: We have implemented the above algorithms in C++ on the Linux operating system.

Algorithms↗

Continuous representations of time-series gene expression data.

We present algorithms for time-series gene expression analysis that permit the principled estimation of unobserved time points, clustering, and dataset alignment. Each expression profile is modeled as a cubic spline (piecewise polynomial) that is estimated from the observed data and every time point influences the overall smooth expression curve. We constrain the spline coefficients of genes in the same class to have similar expression patterns, while also allowing for gene specific parameters. We show that unobserved time points can be reconstructed using our method with 10-15% less error when compared to previous best methods. Our clustering algorithm operates directly on the continuous representations of gene expression profiles, and we demonstrate that this is particularly effective when applied to nonuniformly sampled data. Our continuous alignment algorithm also avoids difficulties encountered by discrete approaches. In particular, our method allows for control of the number of degrees of freedom of the warp through the specification of parameterized functions, which helps to avoid overfitting. We demonstrate that our algorithm produces stable low-error alignments on real expression data and further show a specific application to yeast knock-out data that produces biologically meaningful results.

Algorithms↗