Search PubMed⌕ Search

Biomedical subjects

Gabriela G Loots

Publications and source records attributed to Gabriela G Loots.

16 recordsLinked to original sources

Predicting tissue-specific enhancers in the human genome.

Determining how transcriptional regulatory signals are encoded in vertebrate genomes is essential for understanding the origins of multicellular complexity; yet the genetic code of vertebrate gene regulation remains poorly understood. In an attempt to elucidate this code, we synergistically combined genome-wide gene-expression profiling, vertebrate genome comparisons, and transcription factor binding-site analysis to define sequence signatures characteristic of candidate tissue-specific enhancers in the human genome. We applied this strategy to microarray-based gene expression profiles from 79 human tissues and identified 7187 candidate enhancers that defined their flanking gene expression, the majority of which were located outside of known promoters. We cross-validated this method for its ability to de novo predict tissue-specific gene expression and confirmed its reliability in 57 of the 79 available human tissues, with an average precision in enhancer recognition ranging from 32% to 63% and a sensitivity of 47%. We used the sequence signatures identified by this approach to successfully assign tissue-specific predictions to approximately 328,000 human-mouse conserved noncoding elements in the human genome. By overlapping these genome-wide predictions with a data set of enhancers validated in vivo, in transgenic mice, we were able to confirm our results with a 28% sensitivity and 50% precision. These results indicate the power of combining complementary genomic data sets as an initial computational foray into a global view of tissue-specific gene regulation in vertebrates.

Animals↗

Array2BIO: from microarray expression data to functional annotation of co-regulated genes.

BACKGROUND: There are several isolated tools for partial analysis of microarray expression data. To provide an integrative, easy-to-use and automated toolkit for the analysis of Affymetrix microarray expression data we have developed Array2BIO, an application that couples several analytical methods into a single web based utility. RESULTS: Array2BIO converts raw intensities into probe expression values, automatically maps those to genes, and subsequently identifies groups of co-expressed genes using two complementary approaches: (1) comparative analysis of signal versus control and (2) clustering analysis of gene expression across different conditions. The identified genes are assigned to functional categories based on Gene Ontology classification and KEGG protein interaction pathways. Array2BIO reliably handles low-expressor genes and provides a set of statistical methods for quantifying expression levels, including Benjamini-Hochberg and Bonferroni multiple testing corrections. An automated interface with the ECR Browser provides evolutionary conservation analysis for the identified gene loci while the interconnection with Crème allows prediction of gene regulatory elements that underlie observed expression patterns. CONCLUSION: We have developed Array2BIO - a web based tool for rapid comprehensive analysis of Affymetrix microarray expression data, which also allows users to link expression data to Dcode.org comparative genomics tools and integrates a system for translating co-expression data into mechanisms of gene co-regulation. Array2BIO is publicly available at http://array2bio.dcode.org.

Algorithms↗

Modifying yeast artificial chromosomes to generate Cre/LoxP and FLP/FRT site-specific deletions and inversions.

The ability to efficiently and accurately modify genomic DNA through targeted and tissue-specific mutations is an important goal in animal transgenesis. Here we describe how to exploit two systems of homologous recombination, from yeast and bacteria, to engineer yeast artificial chromosomes (YACs) to generate targeted deletions and inversions in vivo, in transgenic animals, and in the presence of DNA-modifying enzymes known as recombinases. Through homologous recombination in yeast, specific recombinogenic sequences are inserted upstream and downstream of a region in the YAC. The sites of integration of these short sequence elements are chosen carefully, such that the YAC is left functionally intact, and this modified transgene represents the wild-type allele. This YAC is subsequently used to generate transgenic animals, which when bred to animals expressing recombinase proteins result in genetic modifications. By expressing recombinase proteins from different tissue-specific promoters, one can mediate site-specific recombination to generate either ubiquitous or tissue-specific deletions or inversion. These modifications can then be carried through the germline or can be studied somatically. A great advantage of this system is the ability to evaluate subtle genetic effects independent of position-effect variegation, and transgene copy number, eliminating the need to examine several independently generated lines of transgenic animals for each genetic variant.

Animals↗

Dcode.org anthology of comparative genomic tools.

Comparative genomics provides the means to demarcate functional regions in anonymous DNA sequences. The successful application of this method to identifying novel genes is currently shifting to deciphering the non-coding encryption of gene regulation across genomes. To facilitate the practical application of comparative sequence analysis to genetics and genomics, we have developed several analytical and visualization tools for the analysis of arbitrary sequences and whole genomes. These tools include two alignment tools, zPicture and Mulan; a phylogenetic shadowing tool, eShadow for identifying lineage- and species-specific functional elements; two evolutionary conserved transcription factor analysis tools, rVista and multiTF; a tool for extracting cis-regulatory modules governing the expression of co-regulated genes, Creme 2.0; and a dynamic portal to multiple vertebrate and invertebrate genome alignments, the ECR Browser. Here, we briefly describe each one of these tools and provide specific examples on their practical applications. All the tools are publicly available at the http://www.dcode.org/ website.

Base Sequence↗

Genomic deletion of a long-range bone enhancer misregulates sclerostin in Van Buchem disease.

Mutations in distant regulatory elements can have a negative impact on human development and health, yet because of the difficulty of detecting these critical sequences, we predominantly focus on coding sequences for diagnostic purposes. We have undertaken a comparative sequence-based approach to characterize a large noncoding region deleted in patients affected by Van Buchem (VB) disease, a severe sclerosing bone dysplasia. Using BAC recombination and transgenesis, we characterized the expression of human sclerostin (SOST) from normal (SOST(wt)) or Van Buchem (SOST(vbDelta) alleles. Only the SOST(wt) allele faithfully expressed high levels of human SOST in the adult bone and had an impact on bone metabolism, consistent with the model that the VB noncoding deletion removes a SOST-specific regulatory element. By exploiting cross-species sequence comparisons with in vitro and in vivo enhancer assays, we were able to identify a candidate enhancer element that drives human SOST expression in osteoblast-like cell lines in vitro and in the skeletal anlage of the embryonic day 14.5 (E14.5) mouse embryo, and discovered a novel function for sclerostin during limb development. Our approach represents a framework for characterizing distant regulatory elements associated with abnormal human phenotypes.

Adaptor Proteins, Signal Transducing↗

Strategies for characterising cis-regulatory elements in Xenopus.

Understanding the cis-regulatory architecture of metazoan organisms is the greatest challenge facing genome biology today. In vertebrate organisms, distinct sequence elements mediate transcriptional regulation and are scattered throughout the genome, either proximal or distal to promoters. The identification of transcriptional enhancers has proven rather difficult by conventional experimental approaches. In the past decade, the rapid generation of genomic sequences for multiple vertebrate organisms, accompanied by sophisticated comparative tools, has facilitated the identification of non-coding evolutionarily conserved regions that may encode cis-regulatory elements. Validating computational predictions and characterising cis-regulatory elements in vivo, however, has been a major bottleneck, mainly because the most commonly used organism for these experiments has been the mouse, and generating transgenic mice or modifying the mouse genome continues to be a labour-intensive, low-throughput, expensive process. This has led to the use of Xenopus, which holds great promise for high-throughput interrogation of putative cis-regulatory elements. In particular, Xenopus tropicalis may become particularly powerful for elucidating regulatory networks, chiefly because it is amenable to genetic manipulations, and its genome is being sequenced.

Animals↗

Mulan: multiple-sequence local alignment and visualization for studying function and evolution.

Multiple-sequence alignment analysis is a powerful approach for understanding phylogenetic relationships, annotating genes, and detecting functional regulatory elements. With a growing number of partly or fully sequenced vertebrate genomes, effective tools for performing multiple comparisons are required to accurately and efficiently assist biological discoveries. Here we introduce Mulan (http://mulan.dcode.org/), a novel method and a network server for comparing multiple draft and finished-quality sequences to identify functional elements conserved over evolutionary time. Mulan brings together several novel algorithms: the TBA multi-aligner program for rapid identification of local sequence conservation, and the multiTF program for detecting evolutionarily conserved transcription factor binding sites in multiple alignments. In addition, Mulan supports two-way communication with the GALA database; alignments of multiple species dynamically generated in GALA can be viewed in Mulan, and conserved transcription factor binding sites identified with Mulan/multiTF can be integrated and overlaid with extensive genome annotation data using GALA. Local multiple alignments computed by Mulan ensure reliable representation of short- and large-scale genomic rearrangements in distant organisms. Mulan allows for interactive modification of critical conservation parameters to differentially predict conserved regions in comparisons of both closely and distantly related species. We illustrate the uses and applications of the Mulan tool through multispecies comparisons of the GATA3 gene locus and the identification of elements that are conserved in a different way in avians than in other genomes, allowing speculation on the evolution of birds. Source code for the aligners and the aligner-evaluation software can be freely downloaded from http://www.bx.psu.edu/miller_lab/.

Animals↗

Evolution and functional classification of vertebrate gene deserts.

Large tracts of the human genome, known as gene deserts, are devoid of protein-coding genes. Dichotomy in their level of conservation with chicken separates these regions into two distinct categories, stable and variable. The separation is not caused by differences in rates of neutral evolution but instead appears to be related to different biological functions of stable and variable gene deserts in the human genome. Gene Ontology categories of the adjacent genes are strongly biased toward transcriptional regulation and development for the stable gene deserts, and toward distinctively different functions for the variable gene deserts. Stable gene deserts resist chromosomal rearrangements and appear to harbor multiple distant regulatory elements physically linked to their neighboring genes, with the linearity of conservation invariant throughout vertebrate evolution.

Animals↗

ECR Browser: a tool for visualizing and accessing data from comparisons of multiple vertebrate genomes.

With an increasing number of vertebrate genomes being sequenced in draft or finished form, unique opportunities for decoding the language of DNA sequence through comparative genome alignments have arisen. However, novel tools and strategies are required to accommodate this large volume of genomic information and to facilitate the transfer of predictions generated by comparative sequence alignment to researchers focused on experimental annotation of genome function. Here, we present the ECR Browser, a tool that provides easy and dynamic access to whole genome alignments of human, mouse, rat and fish sequences. This web-based tool (http://ecrbrowser.dcode.org) provides the starting point for discovery of novel genes, identification of distant gene regulatory elements and prediction of transcription factor binding sites. The genome alignment portal of the ECR Browser also permits fast and automated alignments of any user-submitted sequence to the genome of choice. The interconnection of the ECR Browser with other DNA sequence analysis tools creates a unique portal for studying and exploring vertebrate genomes.

Animals↗

rVISTA 2.0: evolutionary analysis of transcription factor binding sites.

Identifying and characterizing the transcription factor binding site (TFBS) patterns of cis-regulatory elements represents a challenge, but holds promise to reveal the regulatory language the genome uses to dictate transcriptional dynamics. Several studies have demonstrated that regulatory modules are under positive selection and, therefore, are often conserved between related species. Using this evolutionary principle, we have created a comparative tool, rVISTA, for analyzing the regulatory potential of noncoding sequences. Our ability to experimentally identify functional noncoding sequences is extremely limited, therefore, rVISTA attempts to fill this great gap in genomic analysis by offering a powerful approach for eliminating TFBSs least likely to be biologically relevant. The rVISTA tool combines TFBS predictions, sequence comparisons and cluster analysis to identify noncoding DNA regions that are evolutionarily conserved and present in a specific configuration within genomic sequences. Here, we present the newly developed version 2.0 of the rVISTA tool, which can process alignments generated by both the zPicture and blastz alignment programs or use pre-computed pairwise alignments of several vertebrate genomes available from the ECR Browser and GALA database. The rVISTA web server is closely interconnected with the TRANSFAC database, allowing users to either search for matrices present in the TRANSFAC library collection or search for user-defined consensus sequences. The rVISTA tool is publicly available at http://rvista.dcode.org/.

Algorithms↗

CREME: Cis-Regulatory Module Explorer for the human genome.

The binding of transcription factors to specific regulatory sequence elements is a primary mechanism for controlling gene transcription. Eukaryotic genes are often regulated by several transcription factors whose binding sites are tightly clustered and form cis-regulatory modules. In this paper, we present a web server, CREME, for identifying and visualizing cis-regulatory modules in the promoter regions of a given set of potentially co-regulated genes. CREME relies on a database of putative transcription factor binding sites that have been annotated across the human genome using a library of position weight matrices and evolutionary conservation with the mouse and rat genomes. A search algorithm is applied to this data set to identify combinations of transcription factors whose binding sites tend to co-occur in close proximity in the promoter regions of the input gene set. The identified cis-regulatory modules are statistically scored and significant combinations are reported and graphically visualized. Our web server is available at http://creme.dcode.org.

Binding Sites↗

Interpreting mammalian evolution using Fugu genome comparisons.

Recently, it has been shown that a significant number of evolutionarily conserved human-Fugu noncoding elements function as tissue-specific transcriptional enhancers in vivo, suggesting that distant comparisons are capable of identifying a particular class of regulatory elements. We therefore hypothesized that by juxtaposing human/Fugu and human/mouse conservation patterns we can define conservation criteria for discovering transcriptional regulatory elements specific to mammals. Genome-scale comparisons of noncoding human/Fugu evolutionary conserved elements (ECRs) and their humans/mouse counterparts revealed a particular signature common to human/mouse ECRs (>or=350 bp long, >or=77% identity) that are also conserved in fishes. This newly defined threshold identifies 90% of all human/Fugu noncoding ECRs without the assistance of human-Fugu genome alignments and provides a very efficient filter for identifying functional human/mouse ECRs.

Animals↗

eShadow: a tool for comparing closely related sequences.

Primate sequence comparisons are difficult to interpret due to the high degree of sequence similarity shared between such closely related species. Recently, a novel method, phylogenetic shadowing, has been pioneered for predicting functional elements in the human genome through the analysis of multiple primate sequence alignments. We have expanded this theoretical approach to create a computational tool, eShadow, for the identification of elements under selective pressure in multiple sequence alignments of closely related genomes, such as in comparisons of human-to-primate or mouse-to-rat DNA. This tool integrates two different statistical methods and allows for the dynamic visualization of the resulting conservation profile. eShadow also includes a versatile optimization module capable of training the underlying Hidden Markov Model to differentially predict functional sequences. This module grants the tool high flexibility in the analysis of multiple sequence alignments and in comparing sequences with different divergence rates. Here, we describe the eShadow comparative tool and its potential uses for analyzing both multiple nucleotide and protein alignments to predict putative functional elements.

Animals↗

zPicture: dynamic alignment and visualization tool for analyzing conservation profiles.

Comparative sequence analysis has evolved as an essential technique for identifying functional coding and noncoding elements conserved throughout evolution. Here, we introduce zPicture, an interactive Web-based sequence alignment and visualization tool for dynamically generating conservation profiles and identifying evolutionarily conserved regions (ECRs). zPicture is highly flexible, because critical parameters can be modified interactively, allowing users to differentially predict ECRs in comparisons of sequences of different phylogenetic distances and evolutionary rates. We demonstrate the application of this module to identify a known regulatory element in the HOXD locus, in which functional ECRs are difficult to discern against the highly conserved genomic background. zPicture also facilitates transcription factor binding-site analysis via the rVista tool portal. We present an example of the HBB complex when zPicture/rVista combination specifically pinpoints to two ECRs containing GATA-1, NF-E2, and TAL1/E47 binding sites that were identified previously as transcriptional enhancers. In addition, zPicture is linked to the UCSC Genome Browser, allowing users to automatically extract sequences and gene annotations for any recorded locus. Finally, we describe how this tool can be efficiently applied to the analysis of nonvertebrate genomes, including those of microbial organisms.

Animals↗

rVista for comparative sequence-based discovery of functional transcription factor binding sites.

Identifying transcriptional regulatory elements represents a significant challenge in annotating the genomes of higher vertebrates. We have developed a computational tool, rVista, for high-throughput discovery of cis-regulatory elements that combines clustering of predicted transcription factor binding sites (TFBSs) and the analysis of interspecies sequence conservation to maximize the identification of functional sites. To assess the ability of rVista to discover true positive TFBSs while minimizing the prediction of false positives, we analyzed the distribution of several TFBSs across 1 Mb of the well-annotated cytokine gene cluster (Hs5q31; Mm11). Because a large number of AP-1, NFAT, and GATA-3 sites have been experimentally identified in this interval, we focused our analysis on the distribution of all binding sites specific for these transcription factors. The exploitation of the orthologous human-mouse dataset resulted in the elimination of > 95% of the approximately 58,000 binding sites predicted on analysis of the human sequence alone, whereas it identified 88% of the experimentally verified binding sites in this region.

Animals↗