Search PubMed⌕ Search

Biomedical subjects

M Madan Babu

Publications and source records attributed to M Madan Babu.

At least 19 recordsLinked to original sources

Estimating the prevalence and regulatory potential of the telomere looping effect in yeast transcription regulation.

Telomeres have long been implicated in the regulation of gene expression. Some studies have reported that telomere looping effect (TLE) can juxtapose genes and regulatory sequences that are far apart and facilitate long-distance control of gene expression. In this work, we report a detailed investigation on the prevalence and regulatory potential of TLE on a genomic scale by assembling data on protein-DNA interactions from several large-scale ChIp-chip experiments in Saccharomyces cerevisiae. Analysis of the assembled data revealed that a statistically significant number of DNA segments that were inferred to be bound by ten or more transcription factors in these experiments physically mapped to the ends of several chromosomes (19 of 32 chromosome ends). For the 83 transcription factors that were inferred to interact with these DNA segments, we found a statistically significant skew in the distribution of their internal binding sites over the length of the entire chromosome, such that more than expected binding events occurred proximal to chromosomal ends than elsewhere. Taken together these observations suggest that the telomere looping effect is their most likely explanation and imply that a notable fraction of the internally bound yeast transcription factors potentially interact with looped back telomeres. Further, we also identified several components of the basal transcriptional machinery that are also frequently linked to these chromosome end segments, strengthening the proposal for a direct interaction between the chromosome ends and internally located transcriptional complexes. We observed that certain chromatin factors might participate in the TLE and potentially modulate gene expression by chromatin modifications such as histone deacetylation. Our findings provide the first computational evidence for a significant role of long-range regulatory interactions due to telomere looping. Based on these observations, we also propose that genome-wide chromatin immunoprecipitation data might be useful to systematically uncover long-range chromatin looping effects in gene expression.

Binding Sites↗

Exploring the environmental preference of weak interactions in (alpha/beta)8 barrel proteins.

The environmental preference for the occurrence of noncanonical hydrogen bonding and cation-pi interactions, in a data set containing 71 nonredundant (alpha/beta)(8) barrel proteins, with respect to amino acid type, secondary structure, solvent accessibility, and stabilizing residues has been performed. Our analysis reveals some important findings, which include (a) higher contribution of weak interactions mediated by main-chain atoms irrespective of the amino acids involved; (b) domination of the aromatic amino acids among interactions involving side-chain atoms; (c) involvement of strands as the principal secondary structural unit, accommodating cross strand ion pair interaction and clustering of aromatic amino acid residues; (d) significant contribution to weak interactions occur in the solvent exposed areas of the protein; (e) majority of the interactions involve long-range contacts; (f) the preference of Arg is higher than Lys to form cation-pi interaction; and (g) probability of theoretically predicted stabilizing amino acid residues involved in weak interaction is higher for polar amino acids such as Trp, Glu, and Gln. On the whole, the present study reveals that the weak interactions contribute to the global stability of (alpha/beta)(8) TIM-barrel proteins in an environment-specific manner, which can possibly be exploited for protein engineering applications.

Amino Acids↗

Comprehensive analysis of combinatorial regulation using the transcriptional regulatory network of yeast.

Studies on various model systems have shown that a relatively small number of transcription factors can set up strikingly complex spatial and temporal patterns of gene expression. This is achieved mainly by means of combinatorial or differential gene regulation, i.e. regulation of a gene by two or more transcription factors simultaneously or under different conditions. While a number of specific molecular details of the mechanisms of combinatorial regulation have emerged, our understanding of the general principles of combinatorial regulation on a genomic scale is still limited. In this work, we approach this problem by using the largest assembled transcriptional regulatory network for yeast. A specific network transformation procedure was used to obtain the co-regulatory network describing the set of all significant associations among transcription factors in regulating common target genes. Analysis of the global properties of the co-regulatory network suggested the presence of two classes of regulatory hubs: (i) those that make many co-regulatory associations, thus serving as integrators of disparate cellular processes; and (ii) those that make few co-regulatory associations, and thereby specifically regulate one or a few major cellular processes. Investigation of the local structure of the co-regulatory network revealed a significantly higher than expected modular organization, which might have emerged as a result of selection by functional constraints. These constraints probably emerge from the need for extensive modular backup and the requirement to integrate transcriptional inputs of multiple distinct functional systems. We then explored the transcriptional control of three major regulatory systems (ubiquitin signaling, protein kinase and transcriptional regulation systems) to understand specific aspects of their upstream control. As a result, we observed that ubiquitin E3 ligases are regulated primarily by unique transcription factors, whereas E1 and E2 enzymes share common transcription factors to a much greater extent. This suggested that the deployment of E3s unique to specific functional contexts may be mediated significantly at the transcriptional level. Likewise, we were able to uncover evidence for much higher upstream transcription control of transcription factors themselves, in comparison to components of other regulatory systems. We believe that the results presented here might provide a framework for testing the role of co-regulatory associations in eukaryotic transcriptional control.

Algorithms↗

Uncovering a hidden distributed architecture behind scale-free transcriptional regulatory networks.

Numerous studies in both prokaryotes and eukaryotes have shown that, under standard growth conditions, less than 20% of the protein-coding genes are essential for survival. This suggests that biological systems have evolved to have a high degree of robustness to mutational disruptions that can affect the majority of their genes. This mutational robustness could arise either due to redundancy, i.e. direct backup, or due to distributed architecture, i.e. indirect backup where multiple genes contribute to the functioning of a process in the system. Despite clear evidence for direct backup, the prevalence of indirect backup is poorly understood. In this study, we reveal the existence of a hidden distributed architecture behind the scale-free transcriptional regulatory network of yeast by applying a unique network transformation procedure and show that the network is tolerant even to mutations that disrupt regulatory hubs. Contrary to what is generally accepted, our observation that hubs can be lost or replaced in evolution suggests that this hidden distributed architecture behind scale-free networks protects the overall transcriptional program of the organism from mutations affecting major regulatory hubs. We show that the distributed architecture has been provided by an unexpectedly large number of coordinating partners for any regulatory protein. On the basis of these findings, we propose that the existence of such architecture can allow organisms to explore the adaptive landscape in changing environments by providing the plasticity required to reprogram levels of expression of specific genes that may enhance survival. Thus, an "over-engineered" backup system in the form of distributed architecture is likely to be a major determinant of the "evolvability" of the gene expression in organisms faced with environmental diversity.

Algorithms↗

The HIRAN domain and recruitment of chromatin remodeling and repair activities to damaged DNA.

Aided by sensitive sequence profile searches we identify a novel conserved domain in the N-terminal regions of the SWI2/SNF2 proteins typified by HIP116 and Rad5p (hence HIP116, Rad5p N-terminal domain: HIRAN domain). We show that the HIRAN domain is found as a standalone protein in several bacteria and prophages, or fused to other catalytic domains, such as a nuclease of the restriction endonuclease fold and TDP1-like DNA phosphoesterases, in the eukaryotes. Based on a network of contextual connections in the form of domain architectures, conserved gene neighborhoods and functional interactions we predict that the HIRAN domain is likely to function as a DNA-binding domain that probably recognizes features associated with damaged DNA or stalled replication forks. It might thus act as a sensor to initiate a damaged DNA checkpoint and engage different DNA repair and chromatin remodeling or modifying activities to these sites. In evolutionary terms, the fusion of the HIRAN domain, and the functionally analogous RAD18 Zn-finger and the PARP-type Zn-finger to SWI2/SNF2 ATPases appears to have been a notable factor for recruiting these ATPases for chromatin modification and remodeling in the context of DNA repair.

Amino Acid Sequence↗

Evolutionary dynamics of prokaryotic transcriptional regulatory networks.

The structure of complex transcriptional regulatory networks has been studied extensively in certain model organisms. However, the evolutionary dynamics of these networks across organisms, which would reveal important principles of adaptive regulatory changes, are poorly understood. We use the known transcriptional regulatory network of Escherichia coli to analyse the conservation patterns of this network across 175 prokaryotic genomes, and predict components of the regulatory networks for these organisms. We observe that transcription factors are typically less conserved than their target genes and evolve independently of them, with different organisms evolving distinct repertoires of transcription factors responding to specific signals. We show that prokaryotic transcriptional regulatory networks have evolved principally through widespread tinkering of transcriptional interactions at the local level by embedding orthologous genes in different types of regulatory motifs. Different transcription factors have emerged independently as dominant regulatory hubs in various organisms, suggesting that they have convergently acquired similar network structures approximating a scale-free topology. We note that organisms with similar lifestyles across a wide phylogenetic range tend to conserve equivalent interactions and network motifs. Thus, organism-specific optimal network designs appear to have evolved due to selection for specific transcription factors and transcriptional interactions, allowing responses to prevalent environmental stimuli. The methods for biological network analysis introduced here can be applied generally to study other networks, and these predictions can be used to guide specific experiments.

Amino Acid Motifs↗

Predicting the strongest domain-domain contact in interacting protein pairs.

Experiments to determine the complete 3-dimensional structures of protein complexes are difficult to perform and only a limited range of such structures are available. In contrast, large-scale screening experiments have identified thousands of pairwise interactions between proteins, but such experiments do not produce explicit structural information. In addition, the data produced by these high through-put experiments contain large numbers of false positive results, and can be biased against detection of certain types of interaction. Several methods exist that analyse such pairwise interaction data in terms of the constituent domains within proteins, scoring pairs of domain superfamilies according to their propensity to interact. These scores can be used to predict the strongest domain-domain contact (the contact with the largest surface area) between interacting proteins for which the domain-level structures of the individual proteins are known. We test this predictive approach on a set of pairwise protein interactions taken from the Protein Quaternary Structure (PQS) database for which the true domain-domain contacts are known.While the overall prediction success rate across the whole test data set is poor, we shown how interactions in the test data set for which the training data are not informative can be automatically excluded from the prediction process, giving improved prediction success rates at the expense of restricted coverage of the test data.

Binding Sites↗

A database of bacterial lipoproteins (DOLOP) with functional assignments to predicted lipoproteins.

Lipid modification of the N-terminal Cys residue (N-acyl-S-diacylglyceryl-Cys) has been found to be an essential, ubiquitous, and unique bacterial posttranslational modification. Such a modification allows anchoring of even highly hydrophilic proteins to the membrane which carry out a variety of functions important for bacteria, including pathogenesis. Hence, being able to identify such proteins is of great value. To this end, we have created a comprehensive database of bacterial lipoproteins, called DOLOP, which contains information and links to molecular details for about 278 distinct lipoproteins and predicted lipoproteins from 234 completely sequenced bacterial genomes. The website also features a tool that applies a predictive algorithm to identify the presence or absence of the lipoprotein signal sequence in a user-given sequence. The experimentally verified lipoproteins have been classified into different functional classes and more importantly functional domain assignments using hidden Markov models from the SUPERFAMILY database that have been provided for the predicted lipoproteins. Other features include the following: primary sequence analysis, signal sequence analysis, and search facility and information exchange facility to allow researchers to exchange results on newly characterized lipoproteins. The website, along with additional information on the biosynthetic pathway, statistics on predicted lipoproteins, and related figures, is available at http://www.mrc-lmb.cam.ac.uk/genomes/dolop/.

Bacterial Proteins↗

Adaptive evolution by optimizing expression levels in different environments.

Organisms adapt to environmental changes through the fixation of mutations that enhance reproductive success. A recent study by Dekel and Alon demonstrated that Escherichia coli adapts to different growth conditions by fine-tuning protein levels, as predicted by a simple cost-benefit model. A study by Fong et al. showed that independent evolutionary trajectories lead to similar adaptive endpoints. Initial mutations on the path to adaptation altered the mRNA levels of numerous genes. Subsequent optimization through compensatory mutations restored the expression of most genes to baseline levels, except for a small set that retained differential levels of expression. These studies clarify how adaptation could occur by the alteration of gene expression.

Bacteria↗

The transcriptional landscape of the mammalian genome.

This study describes comprehensive polling of transcription start and termination sites and analysis of previously unidentified full-length complementary DNAs derived from the mouse genome. We identify the 5' and 3' boundaries of 181,047 transcripts with extensive variation in transcripts arising from alternative promoter usage, splicing, and polyadenylation. There are 16,247 new mouse protein-coding transcripts, including 5154 encoding previously unidentified proteins. Genomic mapping of the transcriptome reveals transcriptional forests, with overlapping transcription on both strands, separated by deserts in which few transcripts are observed. The data provide a comprehensive platform for the comparative analysis of mammalian transcriptional regulation in differentiation and development.

3' Untranslated Regions↗

Discovery of the principal specific transcription factors of Apicomplexa and their implication for the evolution of the AP2-integrase DNA binding domains.

The comparative genomics of apicomplexans, such as the malarial parasite Plasmodium, the cattle parasite Theileria and the emerging human parasite Cryptosporidium, have suggested an unexpected paucity of specific transcription factors (TFs) with DNA binding domains that are closely related to those found in the major families of TFs from other eukaryotes. This apparent lack of specific TFs is paradoxical, given that the apicomplexans show a complex developmental cycle in one or more hosts and a reproducible pattern of differential gene expression in course of this cycle. Using sensitive sequence profile searches, we show that the apicomplexans possess a lineage-specific expansion of a novel family of proteins with a version of the AP2 (Apetala2)-integrase DNA binding domain, which is present in numerous plant TFs. About 20-27 members of this apicomplexan AP2 (ApiAP2) family are encoded in different apicomplexan genomes, with each protein containing one to four copies of the AP2 DNA binding domain. Using gene expression data from Plasmodium falciparum, we show that guilds of ApiAP2 genes are expressed in different stages of intraerythrocytic development. By analogy to the plant AP2 proteins and based on the expression patterns, we predict that the ApiAP2 proteins are likely to function as previously unknown specific TFs in the apicomplexans and regulate the progression of their developmental cycle. In addition to the ApiAP2 family, we also identified two other novel families of AP2 DNA binding domains in bacteria and transposons. Using structure similarity searches, we also identified divergent versions of the AP2-integrase DNA binding domain fold in the DNA binding region of the PI-SceI homing endonuclease and the C-terminal domain of the pleckstrin homology (PH) domain-like modules of eukaryotes. Integrating these findings, we present a reconstruction of the evolutionary scenario of the AP2-integrase DNA binding domain fold, which suggests that it underwent multiple independent combinations with different types of mobile endonucleases or recombinases. It appears that the eukaryotic versions have emerged from versions of the domain associated with mobile elements, followed by independent lineage-specific expansions, which accompanied their recruitment to transcription regulation functions.

Amino Acid Sequence↗

The genome of the social amoeba Dictyostelium discoideum.

The social amoebae are exceptional in their ability to alternate between unicellular and multicellular forms. Here we describe the genome of the best-studied member of this group, Dictyostelium discoideum. The gene-dense chromosomes of this organism encode approximately 12,500 predicted proteins, a high proportion of which have long, repetitive amino acid tracts. There are many genes for polyketide synthases and ABC transporters, suggesting an extensive secondary metabolism for producing and exporting small molecules. The genome is rich in complex repeats, one class of which is clustered and may serve as centromeres. Partial copies of the extrachromosomal ribosomal DNA (rDNA) element are found at the ends of each chromosome, suggesting a novel telomere structure and the use of a common mechanism to maintain both the rDNA and chromosomal termini. A proteome-based phylogeny shows that the amoebozoa diverged from the animal-fungal lineage after the plant-animal split, but Dictyostelium seems to have retained more of the diversity of the ancestral genome than have plants, animals or fungi.

ATP-Binding Cassette Transporters↗

Statistical analysis of domains in interacting protein pairs.

MOTIVATION: Several methods have recently been developed to analyse large-scale sets of physical interactions between proteins in terms of physical contacts between the constituent domains, often with a view to predicting new pairwise interactions. Our aim is to combine genomic interaction data, in which domain-domain contacts are not explicitly reported, with the domain-level structure of individual proteins, in order to learn about the structure of interacting protein pairs. Our approach is driven by the need to assess the evidence for physical contacts between domains in a statistically rigorous way. RESULTS: We develop a statistical approach that assigns p-values to pairs of domain superfamilies, measuring the strength of evidence within a set of protein interactions that domains from these superfamilies form contacts. A set of p-values is calculated for SCOP superfamily pairs, based on a pooled data set of interactions from yeast. These p-values can be used to predict which domains come into contact in an interacting protein pair. This predictive scheme is tested against protein complexes in the Protein Quaternary Structure (PQS) database, and is used to predict domain-domain contacts within 705 interacting protein pairs taken from our pooled data set.

Algorithms↗

Evolving nature of the AP2 alpha-appendage hub during clathrin-coated vesicle endocytosis.

Clathrin-mediated endocytosis involves the assembly of a network of proteins that select cargo, modify membrane shape and drive invagination, vesicle scission and uncoating. This network is initially assembled around adaptor protein (AP) appendage domains, which are protein interaction hubs. Using crystallography, we show that FxDxF and WVxF peptide motifs from synaptojanin bind to distinct subdomains on alpha-appendages, called 'top' and 'side' sites. Appendages use both these sites to interact with their binding partners in vitro and in vivo. Occupation of both sites simultaneously results in high-affinity reversible interactions with lone appendages (e.g. eps15 and epsin1). Proteins with multiple copies of only one type of motif bind multiple appendages and so will aid adaptor clustering. These clustered alpha(appendage)-hubs have altered properties where they can sample many different binding partners, which in turn can interact with each other and indirectly with clathrin. In the final coated vesicle, most appendage binding partners are absent and thus the functional status of the appendage domain as an interaction hub is temporal and transitory giving directionality to vesicle assembly.

Adaptor Protein Complex 2↗

Genomic analysis of regulatory network dynamics reveals large topological changes.

Network analysis has been applied widely, providing a unifying language to describe disparate systems ranging from social interactions to power grids. It has recently been used in molecular biology, but so far the resulting networks have only been analysed statically. Here we present the dynamics of a biological network on a genomic scale, by integrating transcriptional regulatory information and gene-expression data for multiple conditions in Saccharomyces cerevisiae. We develop an approach for the statistical analysis of network dynamics, called SANDY, combining well-known global topological measures, local motifs and newly derived statistics. We uncover large changes in underlying network architecture that are unexpected given current viewpoints and random simulations. In response to diverse stimuli, transcription factors alter their interactions to varying degrees, thereby rewiring the network. A few transcription factors serve as permanent hubs, but most act transiently only during certain conditions. By studying sub-network structures, we show that environmental responses facilitate fast signal propagation (for example, with short regulatory cascades), whereas the cell cycle and sporulation direct temporal progression through multiple stages (for example, with highly inter-connected transcription factors). Indeed, to drive the latter processes forward, phase-specific transcription factors inter-regulate serially, and ubiquitously active transcription factors layer above them in a two-tiered hierarchy. We anticipate that many of the concepts presented here--particularly the large-scale topological changes and hub transience--will apply to other biological networks, including complex sub-systems in higher eukaryotes.

Algorithms↗

Gene regulatory network growth by duplication.

We are beginning to elucidate transcriptional regulatory networks on a large scale and to understand some of the structural principles of these networks, but the evolutionary mechanisms that form these networks are still mostly unknown. Here we investigate the role of gene duplication in network evolution. Gene duplication is the driving force for creating new genes in genomes: at least 50% of prokaryotic genes and over 90% of eukaryotic genes are products of gene duplication. The transcriptional interactions in regulatory networks consist of multiple components, and duplication processes that generate new interactions would need to be more complex. We define possible duplication scenarios and show that they formed the regulatory networks of the prokaryote Escherichia coli and the eukaryote Saccharomyces cerevisiae. Gene duplication has had a key role in network evolution: more than one-third of known regulatory interactions were inherited from the ancestral transcription factor or target gene after duplication, and roughly one-half of the interactions were gained during divergence after duplication. In addition, we conclude that evolution has been incremental, rather than making entire regulatory circuits or motifs by duplication with inheritance of interactions.

Bacterial Proteins↗

Structure and evolution of transcriptional regulatory networks.

The regulatory interactions between transcription factors and their target genes can be conceptualised as a directed graph. At a global level, these regulatory networks display a scale-free topology, indicating the presence of regulatory hubs. At a local level, substructures such as motifs and modules can be discerned in these networks. Despite the general organisational similarity of networks across the phylogenetic spectrum, there are interesting qualitative differences among the network components, such as the transcription factors. Although the DNA-binding domains of the transcription factors encoded by a given organism are drawn from a small set of ancient conserved superfamilies, their relative abundance often shows dramatic variation among different phylogenetic groups. Large portions of these networks appear to have evolved through extensive duplication of transcription factors and targets, often with inheritance of regulatory interactions from the ancestral gene. Interactions are conserved to varying degrees among genomes. Insights from the structure and evolution of these networks can be translated into predictions and used for engineering of the regulatory networks of different organisms.

Amino Acid Motifs↗

GenCompass: a universal system for analysing gene expression for any genome.

Microarrays have become indispensable tools for studying the gene expression of particular organisms on a genomic scale. However, despite its widespread use, there are several draw-backs to the current technology. First, it requires prior knowledge of the DNA sequence encoded in the organism of interest, and second, chips must be designed specifically for each genome, greatly increasing the initial cost incurred in manufacturing the arrays.

Chromosome Mapping↗