Search PubMed⌕ Search

Biomedical subjects

Sridhar Hannenhalli

Publications and source records attributed to Sridhar Hannenhalli.

At least 19 recordsLinked to original sources

GATA and Nkx factors synergistically regulate tissue-specific gene expression and development in vivo.

In vitro studies have suggested that members of the GATA and Nkx transcription factor families physically interact, and synergistically activate pulmonary epithelial- and cardiac-gene promoters. However, the relevance of this synergy has not been demonstrated in vivo. We show that Gata6-Titf1 (Gata6-Nkx2.1) double heterozygous (G6-Nkx DH) embryos and mice have severe defects in pulmonary epithelial differentiation and distal airway development, as well as reduced phospholipid production. The defects in G6-Nkx DH embryos and mice are similar to those observed in human neonates with respiratory distress syndromes, including bronchopulmonary dysplasia, and differential gene expression analysis reveals essential developmental pathways requiring synergistic regulation by both Gata6 and Titf1 (Nkx2.1). These studies indicate that Gata6 and Nkx2.1 act in a synergistic manner to direct pulmonary epithelial differentiation and development in vivo, providing direct evidence that interactions between these two transcription factor families are crucial for the development of the tissues in which they are co-expressed.

Alleles↗

Selection of target sites for mobile DNA integration in the human genome.

DNA sequences from retroviruses, retrotransposons, DNA transposons, and parvoviruses can all become integrated into the human genome. Accumulation of such sequences accounts for at least 40% of our genome today. These integrating elements are also of interest as gene-delivery vectors for human gene therapy. Here we present a comprehensive bioinformatic analysis of integration targeting by HIV, MLV, ASLV, SFV, L1, SB, and AAV. We used a mathematical method which allowed annotation of each base pair in the human genome for its likelihood of hosting an integration event by each type of element, taking advantage of more than 200 types of genomic annotation. This bioinformatic resource documents a wealth of new associations between genomic features and integration targeting. The study also revealed that the length of genomic intervals analyzed strongly affected the conclusions drawn--thus, answering the question "What genomic features affect integration?" requires carefully specifying the length scale of interest.

Binding Sites↗

Recurring genomic breaks in independent lineages support genomic fragility.

BACKGROUND: Recent findings indicate that evolutionary breaks in the genome are not randomly distributed, and that certain regions, so-called fragile regions, are predisposed to breakages. Previous approaches to the study of genomic fragility have examined the distribution of breaks, as well as the coincidence of breaks with segmental duplications and repeats, within a single species. In contrast, we investigate whether this regional fragility is an inherent genomic characteristic and is thus conserved over multiple independent lineages. RESULTS: We do this by quantifying the extent to which certain genomic regions are disrupted repeatedly in independent lineages. Our investigation, based on Human, Chimp, Mouse, Rat, Dog and Chicken, suggests that the propensity of a chromosomal region to break is significantly correlated among independent lineages, even when covariates are considered. Furthermore, the fragile regions are enriched for segmental duplications. CONCLUSION: Based on a novel methodology, our work provides additional support for the existence of fragile regions.

Animals↗

Transcriptional genomics associates FOX transcription factors with human heart failure.

BACKGROUND: Specific transcription factors (TFs) modulate cardiac gene expression in murine models of heart failure, but their relevance in human subjects remains untested. We developed and applied a computational approach called transcriptional genomics to test the hypothesis that a discrete set of cardiac TFs is associated with human heart failure. METHODS AND RESULTS: RNA isolates from failing (n=196) and nonfailing (n=16) human hearts were hybridized with Affymetrix HU133A arrays, and differentially expressed heart failure genes were determined. TF binding sites overrepresented in the -5-kb promoter sequences of these heart failure genes were then determined with the use of public genome sequence databases. Binding sites for TFs identified in murine heart failure models (MEF2, NKX, NF-AT, and GATA) were significantly overrepresented in promoters of human heart failure genes (P<0.002; false discovery rate 2% to 4%). In addition, binding sites for FOX TFs showed substantial overrepresentation in both advanced human and early murine heart failure (P<0.002 and false discovery rate <4% for each). A role for FOX TFs was supported further by expression of FOXC1, C2, P1, P4, and O1A in failing human cardiac myocytes at levels similar to established hypertrophic TFs and by abundant FOXP1 protein in failing human cardiac myocyte nuclei. CONCLUSIONS: Our results provide the first evidence that specific TFs identified in murine models (MEF2, NKX, NFAT, and GATA) are associated with human heart failure. Moreover, these data implicate specific members of the FOX family of TFs (FOXC1, C2, P1, P4, and O1A) not previously suggested in heart failure pathogenesis. These findings provide a crucial link between animal models and human disease and suggest a specific role for FOX signaling in modulating the hypertrophic response of the heart to stress in humans.

Adult↗

Dense subgraph computation via stochastic search: application to detect transcriptional modules.

MOTIVATION: In a tri-partite biological network of transcription factors, their putative target genes, and the tissues in which the target genes are differentially expressed, a tightly inter-connected (dense) subgraph may reveal knowledge about tissue specific transcription regulation mediated by a specific set of transcription factors-a tissue-specific transcriptional module. This is just one context in which an efficient computation of dense subgraphs in a multi-partite graph is needed. RESULT: Here we report a generic stochastic search based method to compute dense subgraphs in a graph with an arbitrary number of partitions and an arbitrary connectivity among the partitions. We then use the tool to explore tissue-specific transcriptional regulation in the human genome. We validate our findings in Skeletal muscle based on literature. We could accurately deduce biological processes for transcription factors via the tri-partite clusters of transcription factors, genes, and the functional annotation of genes. Additionally, we propose a few previously unknown TF-pathway associations and tissue-specific roles for certain pathways. Finally, our combined analysis of Cardiac, Skeletal, and Smooth muscle data recapitulates the evolutionary relationship among the three tissues.

Algorithms↗

Retroviral DNA integration: viral and cellular determinants of target-site selection.

Retroviruses differ in their preferences for sites for viral DNA integration in the chromosomes of infected cells. Human immunodeficiency virus (HIV) integrates preferentially within active transcription units, whereas murine leukemia virus (MLV) integrates preferentially near transcription start sites and CpG islands. We investigated the viral determinants of integration-site selection using HIV chimeras with MLV genes substituted for their HIV counterparts. We found that transferring the MLV integrase (IN) coding region into HIV (to make HIVmIN) caused the hybrid to integrate with a specificity close to that of MLV. Addition of MLV gag (to make HIVmGagmIN) further increased the similarity of target-site selection to that of MLV. A chimeric virus with MLV Gag only (HIVmGag) displayed targeting preferences different from that of both HIV and MLV, further implicating Gag proteins in targeting as well as IN. We also report a genome-wide analysis indicating that MLV, but not HIV, favors integration near DNase I-hypersensitive sites (i.e., +/- 1 kb), and that HIVmIN and HIVmGagmIN also favored integration near these features. These findings reveal that IN is the principal viral determinant of integration specificity; they also reveal a new role for Gag-derived proteins, and strengthen models for integration targeting based on tethering of viral IN proteins to host proteins.

Attachment Sites, Microbiological↗

A mammalian promoter model links cis elements to genetic networks.

An accurate identification of gene promoters remains an important challenge. Computational approaches for this problem rely on promoter sequence attributes that are believed to be critical for transcription initiation. Here we report a probabilistic model that captures two important properties of promoters, not used by previous methods, viz., the location preference and co-occurrence of promoter elements. Additionally, we found that many of the position-specific DNA elements are strongly linked with the function of the gene product. For instance, a highly conserved motif CCTTT at -1 position is strongly associated with protein synthesis, cellular and tissue development. Our comparative analysis of promoter classes reveals that the promoters devoid of CpG islands are more conserved and have fewer alternative transcription start sites. The discovered links between promoter elements and gene function allows us to infer genetic networks from promoter elements. The web server for the PSPA promoter predictor is available at /PSPA.

Animals↗

Differential transcriptional response to nonassociative and associative components of classical fear conditioning in the amygdala and hippocampus.

Classical fear conditioning requires the recognition of conditioned stimuli (CS) and the association of the CS with an aversive stimulus. We used Affymetrix oligonucleotide microarrays to characterize changes in gene expression compared to naive mice in both the amygdala and the hippocampus 30 min after classical fear conditioning and 30 min after exposure to the CS in the absence of an aversive stimulus. We found that in the hippocampus, levels of gene regulation induced by classical fear conditioning were not significantly greater than those induced by CS alone, whereas in the amygdala, classical fear conditioning did induce significantly greater levels of gene regulation compared to the CS. Computational studies suggest that transcriptional changes in the hippocampus and amygdala are mediated by large and overlapping but distinct combinations of molecular events. Our results demonstrate that an increase in gene regulation in the amygdala was partially correlated to associative learning and partially correlated to nonassociative components of the task, while gene regulation in the hippocampus was correlated to nonassociative components of classical fear conditioning, including configural learning.

Amygdala↗

Patterns of sequence conservation in presynaptic neural genes.

BACKGROUND: The neuronal synapse is a fundamental functional unit in the central nervous system of animals. Because synaptic function is evolutionarily conserved, we reasoned that functional sequences of genes and related genomic elements known to play important roles in neurotransmitter release would also be conserved. RESULTS: Evolutionary rate analysis revealed that presynaptic proteins evolve slowly, although some members of large gene families exhibit accelerated evolutionary rates relative to other family members. Comparative sequence analysis of 46 megabases spanning 150 presynaptic genes identified more than 26,000 elements that are highly conserved in eight vertebrate species, as well as a small subset of sequences (6%) that are shared among unrelated presynaptic genes. Analysis of large gene families revealed that upstream and intronic regions of closely related family members are extremely divergent. We also identified 504 exceptionally long conserved elements (> or =360 base pairs, > or =80% pair-wise identity between human and other mammals) in intergenic and intronic regions of presynaptic genes. Many of these elements form a highly stable stem-loop RNA structure and consequently are candidates for novel regulatory elements, whereas some conserved noncoding elements are shown to correlate with specific gene expression profiles. The SynapseDB online database integrates these findings and other functional genomic resources for synaptic genes. CONCLUSION: Highly conserved elements in nonprotein coding regions of 150 presynaptic genes represent sequences that may be involved in the transcriptional or post-transcriptional regulation of these genes. Furthermore, comparative sequence analysis will facilitate selection of genes and noncoding sequences for future functional studies and analysis of variation studies in neurodevelopmental and psychiatric disorders.

Animals↗

Functional analysis of Hes-1 in preadipocytes.

Notch signaling blocks differentiation of 3T3-L1 preadipocytes, and this can be mimicked by constitutive expression of the Notch target gene Hes-1. Although considered initially to function only as a repressor, recent evidence indicates that Hes-1 can also activate transcription. We show here that the domains of Hes-1 needed to block adipogenesis coincide with those necessary for transcriptional repression. HRT1, another basic-helix-loop-helix protein and potential Hes-1 partner, was also induced by Notch in 3T3-L1 cells but did not block adipogenesis, suggesting that Hes-1 functions primarily as a homodimer or possibly as a heterodimer with an unknown partner. Purification of Hes-1 identified the Groucho/transducin-like enhancer of split family of corepressors as the only significant Hes-1 interacting proteins in vivo. An evaluation of global gene expression in preadipocytes identified approximately 200 Hes-1-responsive genes comprising roughly equal numbers of up-regulated and down-regulated genes. However, promoter analyses indicated that the down-regulated genes were significantly more likely to contain Hes-1 binding sites, indicating that Hes-1 is more likely to repress transcription of its direct targets. We conclude that Notch most likely blocks adipogenesis through the induction of Hes-1 homodimers, which repress transcription of key target genes.

Adipocytes↗

Generalizations of Markov model to characterize biological sequences.

BACKGROUND: The currently used kth order Markov models estimate the probability of generating a single nucleotide conditional upon the immediately preceding (gap = 0) k units. However, this neither takes into account the joint dependency of multiple neighboring nucleotides, nor does it consider the long range dependency with gap > 0. RESULT: We describe a configurable tool to explore generalizations of the standard Markov model. We evaluated whether the sequence classification accuracy can be improved by using an alternative set of model parameters. The evaluation was done on four classes of biological sequences--CpG-poor promoters, all promoters, exons and nucleosome positioning sequences. Using di- and tri-nucleotide as the model unit significantly improved the sequence classification accuracy relative to the standard single nucleotide model. In the case of nucleosome positioning sequences, optimal accuracy was achieved at a gap length of 4. Furthermore in the plot of classification accuracy versus the gap, a periodicity of 10-11 bps was observed which might indicate structural preferences in the nucleosome positioning sequence. The tool is implemented in Java and is available for download at ftp://ftp.pcbi.upenn.edu/GMM/. CONCLUSION: Markov modeling is an important component of many sequence analysis tools. We have extended the standard Markov model to incorporate joint and long range dependencies between the sequence elements. The proposed generalizations of the Markov model are likely to improve the overall accuracy of sequence analysis tools.

Base Sequence↗

Genome-wide analysis of retroviral DNA integration.

Retroviral vectors are often used to introduce therapeutic sequences into patients' cells. In recent years, gene therapy with retroviral vectors has had impressive therapeutic successes, but has also resulted in three cases of leukaemia caused by insertional mutagenesis, which has focused attention on the molecular determinants of retroviral-integration target-site selection. Here, we review retroviral DNA integration, with emphasis on recent genome-wide studies of targeting and on the status of efforts to modulate target-site selection.

Animals↗

Enhanced position weight matrices using mixture models.

MOTIVATION: Positional weight matrix (PWM) is derived from a set of experimentally determined binding sites. Here we explore whether there exist subclasses of binding sites and if the mixture of these subclass-PWMs can improve the binding site prediction. Intuitively, the subclasses correspond to either distinct binding preference of the same transcription factor in different contexts or distinct subtypes of the transcription factor. AVAILABILITY: We report an Expectation Maximization algorithm adapting the mixture model of Baily and Elkan. We assessed the relative merit of using two subclass-PWMs. The resulting PWMs were evaluated with respect to preferred conservation (relative to mouse) of potential sites in human promoters and expression coherence of the potential target genes. Based on 64 JASPAR vertebrate PWMs, 61-81% of the cases resulted in a higher conservation using the mixture model. Also in 98% of the cases the expression coherence was higher for the target genes of one of the subclass-PWMs. Our analysis of Reb1 sites is consistent with previously discovered subtypes using independent methods. Additionally application of our method to mutated sites for transcription factor LEU3 reveals subclasses that segregate into strongly binding and weakly binding sites with P-value of 0.008. This is the first study which attempts to quantify the subtly different binding specificities of a transcription factor on a large scale and suggests the use of a mixture of PWMs, instead of the current practice of using a single PWM, for a transcription factor.

Algorithms↗

Whole-genome shotgun assembly and comparison of human genome assemblies.

We report a whole-genome shotgun assembly (called WGSA) of the human genome generated at Celera in 2001. The Celera-generated shotgun data set consisted of 27 million sequencing reads organized in pairs by virtue of end-sequencing 2-kbp, 10-kbp, and 50-kbp inserts from shotgun clone libraries. The quality-trimmed reads covered the genome 5.3 times, and the inserts from which pairs of reads were obtained covered the genome 39 times. With the nearly complete human DNA sequence [National Center for Biotechnology Information (NCBI) Build 34] now available, it is possible to directly assess the quality, accuracy, and completeness of WGSA and of the first reconstructions of the human genome reported in two landmark papers in February 2001 [Venter, J. C., Adams, M. D., Myers, E. W., Li, P. W., Mural, R. J., Sutton, G. G., Smith, H. O., Yandell, M., Evans, C. A., Holt, R. A., et al. (2001) Science 291, 1304-1351; International Human Genome Sequencing Consortium (2001) Nature 409, 860-921]. The analysis of WGSA shows 97% order and orientation agreement with NCBI Build 34, where most of the 3% of sequence out of order is due to scaffold placement problems as opposed to assembly errors within the scaffolds themselves. In addition, WGSA fills some of the remaining gaps in NCBI Build 34. The early genome sequences all covered about the same amount of the genome, but they did so in different ways. The Celera results provide more order and orientation, and the consortium sequence provides better coverage of exact and nearly exact repeats.

Computational Biology↗

A note on efficient computation of haplotypes via perfect phylogeny.

The problem of inferring haplotype phase from a population of genotypes has received a lot of attention recently. This is partly due to the observation that there are many regions on human genomic DNA where genetic recombination is rare (Helmuth, 2001; Daly et al., 2001; Stephens et al., 2001; Friss et al., 2001). A Haplotype Map project has been announced by NIH to identify and characterize populations in terms of these haplotypes. Recently, Gusfield introduced the perfect phylogeny haplotyping problem, as an algorithmic implication of the no-recombination in long blocks observation, together with the standard population-genetic assumption of infinite sites. Gusfield's solution based on matroid theory was followed by direct theta(nm2) solutions that use simpler techniques (Bafna et al., 2003; Eskin et al., 2003), and also bound the number of solutions to the PPH problem. In this short note, we address two questions that were left open. First, can the algorithms of Bafna et al. (2003) and Eskin et al. (2003) be sped-up to O(nm + m2) time, which would imply an O(nm) time-bound for the PPH problem? Second, if there are multiple solutions, can we find one that is most parsimonious in terms of the number of distinct haplotypes. We give reductions that suggests that the answer to both questions is "no." For the first problem, we show that computing the output of the first step (in either method) is equivalent to Boolean matrix multiplication. Therefore, the best bound we can presently achieve is O(nm(omega-1)), where omega < or = 2.52 is the exponent of matrix multiplication. Thus, any linear time solution to the PPH problem likely requires a different approach. For the second problem of computing a PPH solution that minimizes the number of distinct haplotypes, we show that the problem is NP-hard using a reduction from Vertex Cover (Garey and Johnson, 1979).

Computational Biology↗

Transcriptional regulation of protein complexes and biological pathways.

The cis-element profile (or cis-profile) of a gene refers to the collection of transcription factor binding sites (TFBS) regulating the transcription of the gene. Underlying the various published studies that attempt to discover cis-elements in the vicinity of co-expressed genes via pattern detection algorithms, there is an implicit assumption that a correlation exists between co-expressed genes and their cis-profiles. In this study, we show that the cis-similarity, defined as the proportion of shared TFBS between two cis-element profiles, is higher for functionally linked interacting proteins as well as for members of a signal transduction pathway. A similar analysis of the enzymes catalyzing the conversion of adjacent substrates to products in a collection of metabolic pathways, did not reveal higher cis-similarity. The analysis is based on three distinct sources of publicly available data, namely, 1) the BIND database of interacting proteins, 2) known interactions in NMDAR protein complex, 3) the apoptosis pathway and nine pathways related to metabolism of cofactors and vitamins all from KEGG. Additionally, we analyze the cis-element profiles of all the genes in the glutamate receptor (GR) sub-complex of NMDAR complex to detect a set of cis-elements that occur adjacent to a majority of the genes. We show that most of the corresponding transcription factors are known to be involved in GR regulation by comparing our findings with the published biomedical literature. In addition, we were able to detect transcripts whose gene products associate with GR by searching for transcripts that share the same regulatory signals as those detected for GR. This suggests a novel computational methodology for constructing high-order gene regulatory models and detecting co-regulated gene products.

Apoptosis↗

Predicting transcription factor synergism.

Transcriptional regulation is mediated by a battery of transcription factor (TF) proteins, that form complexes involving protein-protein and protein-DNA interactions. Individual TFs bind to their cognate cis-elements or transcription factor-binding sites (TFBS). TFBS are organized on the DNA proximal to the gene in groups confined to a few hundred base pair regions. These groups are referred to as modules. Various modules work together to provide the combinatorial regulation of gene transcription in response to various developmental and environmental conditions. The sets of modules constitute a promoter model. Determining the TFs that preferentially work in concert as part of a module is an essential component of understanding transcriptional regulation. The TFs that act synergistically in such a fashion are likely to have their cis-elements co-localized on the genome at specific distances apart. We exploit this notion to predict TF pairs that are likely to be part of a transcriptional module on the human genome sequence. The computational method is validated statistically, using known interacting pairs extracted from the literature. There are 251 TFBS pairs up to 50 bp apart and 70 TFBS pairs up to 200 bp apart that score higher than any of the known synergistic pairs. Further investigation of 50 pairs randomly selected from each of these two sets using PubMed queries provided additional supporting evidence from the existing biological literature suggesting TF synergism for these novel pairs.

Algorithms↗

Combinatorial algorithms for design of DNA arrays.

Optimal design of DNA arrays requires the development of algorithms with two-fold goals: reducing the effects caused by unintended illumination (border length minimization problem) and reducing the complexity of masks (mask decomposition problem). We describe algorithms that reduce the number of rectangles in mask decomposition by 20-30% as compared to a standard array design under the assumption that the arrangement of oligonucleotides on the array is fixed. This algorithm produces provably optimal solution for all studied real instances of array design. We also address the difficult problem of finding an arrangement which minimizes the border length and come up with a new idea of threading that significantly reduces the border length as compared to standard designs.

Algorithms↗