Search PubMed⌕ Search

Biomedical subjects

L Aravind

Publications and source records attributed to L Aravind.

At least 37 records · Page 2Linked to original sources

The ASCH superfamily: novel domains with a fold related to the PUA domain and a potential role in RNA metabolism.

Several studies show that transcription coactivators are often bi-functional ribonucleoprotein complexes that also regulate pre-mRNA processing and splicing decisions. Using sensitive sequence profile searches and structural comparisons we show that the C-terminal domain of the human coactivator protein ASC-1 defines a novel superfamily, the ASC-1 homology (ASCH) domain. The approximately 110 amino acid long ASCH domains are widely represented in all the three superkingdoms of life and several prokaryotic viruses. We show that the ASCH superfamily adopts a beta-barrel fold similar to the PUA domain superfamily. Using multiple lines of evidence, we suggest that members of the ASCH superfamily are likely to function as RNA-binding domains in contexts related to coactivation, RNA-processing and possibly prokaryotic translation regulation. Structural analysis of ASCH domains reveals the presence of a potential RNA-binding cleft associated with a conserved sequence motif, which is characteristic of this superfamily. Despite their similar structure, the ASCH and PUA domains appear to occupy distinct functional niches, with the former domains typically occurring in a standalone form in polypeptides, and the latter domains showing fusions to a variety of RNA-modifying enzymes.

Animals↗

Discovery of the principal specific transcription factors of Apicomplexa and their implication for the evolution of the AP2-integrase DNA binding domains.

The comparative genomics of apicomplexans, such as the malarial parasite Plasmodium, the cattle parasite Theileria and the emerging human parasite Cryptosporidium, have suggested an unexpected paucity of specific transcription factors (TFs) with DNA binding domains that are closely related to those found in the major families of TFs from other eukaryotes. This apparent lack of specific TFs is paradoxical, given that the apicomplexans show a complex developmental cycle in one or more hosts and a reproducible pattern of differential gene expression in course of this cycle. Using sensitive sequence profile searches, we show that the apicomplexans possess a lineage-specific expansion of a novel family of proteins with a version of the AP2 (Apetala2)-integrase DNA binding domain, which is present in numerous plant TFs. About 20-27 members of this apicomplexan AP2 (ApiAP2) family are encoded in different apicomplexan genomes, with each protein containing one to four copies of the AP2 DNA binding domain. Using gene expression data from Plasmodium falciparum, we show that guilds of ApiAP2 genes are expressed in different stages of intraerythrocytic development. By analogy to the plant AP2 proteins and based on the expression patterns, we predict that the ApiAP2 proteins are likely to function as previously unknown specific TFs in the apicomplexans and regulate the progression of their developmental cycle. In addition to the ApiAP2 family, we also identified two other novel families of AP2 DNA binding domains in bacteria and transposons. Using structure similarity searches, we also identified divergent versions of the AP2-integrase DNA binding domain fold in the DNA binding region of the PI-SceI homing endonuclease and the C-terminal domain of the pleckstrin homology (PH) domain-like modules of eukaryotes. Integrating these findings, we present a reconstruction of the evolutionary scenario of the AP2-integrase DNA binding domain fold, which suggests that it underwent multiple independent combinations with different types of mobile endonucleases or recombinases. It appears that the eukaryotic versions have emerged from versions of the domain associated with mobile elements, followed by independent lineage-specific expansions, which accompanied their recruitment to transcription regulation functions.

Amino Acid Sequence↗

Origin and evolution of the archaeo-eukaryotic primase superfamily and related palm-domain proteins: structural insights and new members.

We report an in-depth computational study of the protein sequences and structures of the superfamily of archaeo-eukaryotic primases (AEPs). This analysis greatly expands the range of diversity of the AEPs and reveals the unique active site shared by all members of this superfamily. In particular, it is shown that eukaryotic nucleo-cytoplasmic large DNA viruses, including poxviruses, asfarviruses, iridoviruses, phycodnaviruses and the mimivirus, encode AEPs of a distinct family, which also includes the herpesvirus primases whose relationship to AEPs has not been recognized previously. Many eukaryotic genomes, including chordates and plants, encode previously uncharacterized homologs of these predicted viral primases, which might be involved in novel DNA repair pathways. At a deeper level of evolutionary connections, structural comparisons indicate that AEPs, the nucleases involved in the initiation of rolling circle replication in plasmids and viruses, and origin-binding domains of papilloma and polyoma viruses evolved from a common ancestral protein that might have been involved in a protein-priming mechanism of initiation of DNA replication. Contextual analysis of multidomain protein architectures and gene neighborhoods in prokaryotes and viruses reveals remarkable parallels between AEPs and the unrelated DnaG-type primases, in particular, tight associations with the same repertoire of helicases. These observations point to a functional equivalence of the two classes of primases, which seem to have repeatedly displaced each other in various extrachromosomal replicons.

Amino Acid Sequence↗

MEDS and PocR are novel domains with a predicted role in sensing simple hydrocarbon derivatives in prokaryotic signal transduction systems.

UNLABELLED: We identify two conserved domains in diverse bacterial and archaeal signaling proteins. One of them, the MEDS domain, is typified by the DmcR protein from Methylococcus and the other by the PocR protein of Salmonella typhi. We provide evidence that both these domains are likely to sense simple hydrocarbon derivatives and transduce downstream signals on binding these ligands. The PocR ligand-binding domain is shown to contain a novel variant of the fold found in PAS and GAF domains. The MEDS domain is present in both methylotrophs and complex methanogens, and both the MEDS and PocR domains show a lineage-specific expansion in the latter organisms, suggesting a role in sensing their principle growth substrates. The MEDS domain is also found in the negative regulators of the sigma factor SigB in actinomycetes, including pathogens like Mycobacterium tuberculosis. Hence it is possible that these sigma factors, involved in aerial mycelium development and stress response in the actinomycetes, might be under the regulation of as yet uncharacterized small molecules. CONTACT: aravind@ncbi.nlm.nih.gov.

Bacterial Proteins↗

Plasmodium falciparum: characterization of a late asexual stage golgi protein containing both ankyrin and DHHC domains.

Proteins containing the DHHC motif have been shown to function as palmitoyl transferases. The palmitoylation of proteins has been shown to play an important role in the trafficking of proteins to the proper subcellular location. Herein, we describe a protein containing both ankyrin domains and a DHHC domain that is present in the Golgi of late schizonts of P. falciparum. The timing of expression as well as the location of this protein suggests that it may play an important role in the sorting of proteins to the apical organelles during the development of the asexual stage of the parasite.

Amino Acid Sequence↗

The many faces of the helix-turn-helix domain: transcription regulation and beyond.

The helix-turn-helix (HTH) domain is a common denominator in basal and specific transcription factors from the three super-kingdoms of life. At its core, the domain comprises of an open tri-helical bundle, which typically binds DNA with the 3rd helix. Drawing on the wealth of data that has accumulated over two decades since the discovery of the domain, we present an overview of the natural history of the HTH domain from the viewpoint of structural analysis and comparative genomics. In structural terms, the HTH domains have developed several elaborations on the basic 3-helical core, such as the tetra-helical bundle, the winged-helix and the ribbon-helix-helix type configurations. In functional terms, the HTH domains are present in the most prevalent transcription factors of all prokaryotic genomes and some eukaryotic genomes. They have been recruited to a wide range of functions beyond transcription regulation, which include DNA repair and replication, RNA metabolism and protein-protein interactions in diverse signaling contexts. Beyond their basic role in mediating macromolecular interactions, the HTH domains have also been incorporated into the catalytic domains of diverse enzymes. We discuss the general domain architectural themes that have arisen amongst the HTH domains as a result of their recruitment to these diverse functions. We present a natural classification, higher-order relationships and phyletic pattern analysis of all the major families of HTH domains. This reconstruction suggests that there were at least 6-11 different HTH domains in the last universal common ancestor of all life forms, which covered much of the structural diversity and part of the functional versatility of the extant representatives of this domain. In prokaryotes the total number of HTH domains per genome shows a strong power-equation type scaling with the gene number per genome. However, the HTH domains in two-component signaling pathways show a linear scaling with gene number, in contrast to the non-linear scaling of HTH domains in single-component systems and sigma factors. These observations point to distinct evolutionary forces in the emergence of different signaling systems with HTH transcription factors. The archaea and bacteria share a number of ancient families of specific HTH transcription factors. However, they do not share any orthologous HTH proteins in the basal transcription apparatus. This differential relationship of their basal and specific transcriptional machinery poses an apparent conundrum regarding the origins of their transcription apparatus.

Amino Acid Sequence↗

Identification of the prokaryotic ligand-gated ion channels and their implications for the mechanisms and origins of animal Cys-loop ion channels.

BACKGROUND: Acetylcholine receptor type ligand-gated ion channels (ART-LGIC; also known as Cys-loop receptors) are a superfamily of proteins that include the receptors for major neurotransmitters such as acetylcholine, serotonin, glycine, GABA, glutamate and histamine, and for Zn2+ ions. They play a central role in fast synaptic signaling in animal nervous systems and so far have not been found outside of the Metazoa. RESULTS: Using sensitive sequence-profile searches we have identified homologs of ART-LGICs in several bacteria and a single archaeal genus, Methanosarcina. The homology between the animal receptors and the prokaryotic homologs spans the entire length of the former, including both the ligand-binding and channel-forming transmembrane domains. A sequence-structure analysis using the structure of Lymnaea stagnalis acetylcholine-binding protein and the newly detected prokaryotic versions indicates the presence of at least one aromatic residue in the ligand-binding boxes of almost all representatives of the superfamily. Investigation of the domain architectures of the bacterial forms shows that they may often show fusions with other small-molecule-binding domains, such as the periplasmic binding protein superfamily I (PBP-I), Cache and MCP-N domains. Some of the bacterial forms also occur in predicted operons with the genes of the PBP-II superfamily and the Cache domains. Analysis of phyletic patterns suggests that the ART-LGICs are currently absent in all other eukaryotic lineages except animals. Moreover, phylogenetic analysis and conserved sequence motifs also suggest that a subset of the bacterial forms is closer to the metazoan forms. CONCLUSIONS: From the information from the bacterial forms we infer that cation-pi or hydrophobic interactions with the ligand are likely to be a pervasive feature of the entire superfamily, even though the individual residues involved in the process may vary. The conservation pattern in the channel-forming transmembrane domains also suggests similar channel-gating mechanisms in the prokaryotic versions. From the distribution of charged residues in the prokaryotic M2 transmembrane segments, we expect that there will be examples of both cation and anion selectivity within the prokaryotic members. Contextual connections suggest that the prokaryotic forms may function as chemotactic receptors for low molecular weight solutes. The phyletic patterns and phylogenetic relationships suggest the possibility that the metazoan receptors emerged through an early lateral transfer from a prokaryotic source, before the divergence of extant metazoan lineages.

Animals↗

Comparative genomics, evolution and origins of the nuclear envelope and nuclear pore complex.

The presence of a distinct nucleus, the compartment for confining the genome, transcription and RNA maturation, is a central (and eponymous) feature that distinguishes eukaryotes from prokaryotes. Structural integrity of the nucleus is maintained by the nuclear envelope (NE). A crucial element of this structure is the nuclear pore complex (NPC), a macromolecular machine with over 90 protein components, which mediates nucleo-cytoplasmic communication. We investigated the provenance of the conserved domains found in these perinuclear proteins and reconstructed a parsimonious scenario for NE and NPC evolution by means of comparative-genomic analysis of their components from the available sequences of 28 sequenced eukaryotic genomes. We show that the NE and NPC proteins were tinkered together from diverse domains, which evolved from prokaryotic precursors at different points in eukaryotic evolution, divergence from pre-existing eukaryotic paralogs performing other functions, and de novo. It is shown that several central components of the NPC, in particular, the RanGDP import factor NTF2, the HEH domain of Src1p-Man1, and, probably, also the key domains of karyopherins and nucleoporins, the HEAT/ARM and WD40 repeats, have a bacterial, most likely, endosymbiotic origin. The specialized immunoglobulin (Ig) domain in the globular tail of the animal lamins, and the Ig domains in the nuclear membrane protein GP210 are shown to be related to distinct prokaryotic families of Ig domains. This suggests that independent, late horizontal gene transfer events from bacterial sources might have contributed to the evolution of perinuclear proteins in some of the major eukaryotic lineages. Snurportin 1, one of the highly conserved karyopherins, contains a cap-binding domain which is shown to be an inactive paralog of the guanylyl transferase domain of the mRNA-capping enzyme, exemplifying recruitment of paralogs of pre-exsiting proteins for perinuclear functions. It is shown that several NPC proteins containing super-structure- forming alpha-helical and beta-propeller modules are most closely related to corresponding proteins in the cytoplasmic vesicle biogenesis and coating complexes. From these observations, we infer an autogenous scenario of nuclear evolution in which the nucleus emerged in the primitive eukaryotic ancestor (the "prekaryote") as part of cell compartmentalization triggered by archaeo-bacterial symbiosis. A pivotal event in this process was the radiation of Ras-superfamily GTPases yielding Ran, the key regulator of nuclear transport. A primitive NPC with approximately 20 proteins and a Src1p-Man1-like membrane protein with a DNA-tethering HEH domain are inferred to have been integral perinuclear components in the las common ancestor of modern eukaryotes.

Amino Acid Sequence↗

Novel predicted peptidases with a potential role in the ubiquitin signaling pathway.

A multi-pronged strategy including extensive sequence searches, structural modeling, and analysis of contextual information extracted from domain architectures, genetic screens, and large-scale protein-protein interaction analyses was employed to predict previously undetected components of the eukaryotic ubiquitin (Ub) signaling system. Two novel groups of proteins that are likely to function as de-ubiquitinating and de-SUMOylating peptidases (DUBs) were identified. The first group of putative DUBs, designated PPPDE superfamily (after Permuted Papain fold Peptidases of DsRNA viruses and Eukaryotes), consists of predicted thiol peptidases with a circularly permuted papain-like fold. The inference of the likely DUB function of the PPPDE superfamily proteins is based on the fusions of the catalytic domain to Ub-binding PUG (PUB)/UBA domains and a novel alpha-helical Ub-associated domain (the PUL domain, after PLAP, Ufd3p and Lub1p). The presence of the PPPDE superfamily proteins in most eukaryotic lineages, including basal ones, such as Giardia, suggests a role in deubiquitination of highly conserved proteins involved in key cellular functions, such as cell cycle control. In addition to eukaryotic proteins, the PPPDE superfamily includes predicted proteases from several groups of double-stranded RNA viruses and one single-stranded DNA virus. The apparent recruitment of DUBs for viral polyprotein processing seems to represent a common theme in evolution of viruses. The second group of putative DUBs identified in this study is the WLM (Wss1p-like metalloproteases) family of the Zincin-like superfamily of Zn-dependent peptidases, which are linked to the Ub-system by virtue of fusions with the UB-binding PUG (PUB), Ub-like, and Little Finger domains. More specifically, genetic evidence implicates the WLM family in de-SUMOylation. If validated experimentally, the WLM family proteins will represent the first case of a Zincin-like metalloprotease involvement in Ub-signaling.

Amino Acid Sequence↗

STAND, a class of P-loop NTPases including animal and plant regulators of programmed cell death: multiple, complex domain architectures, unusual phyletic patterns, and evolution by horizontal gene transfer.

Using sequence profile analysis and sequence-based structure predictions, we define a previously unrecognized, widespread class of P-loop NTPases. The signal transduction ATPases with numerous domains (STAND) class includes the AP-ATPases (animal apoptosis regulators CED4/Apaf-1, plant disease resistance proteins, and bacterial AfsR-like transcription regulators) and NACHT NTPases (e.g. NAIP, TLP1, Het-E-1) that have been studied extensively in the context of apoptosis, pathogen response in animals and plants, and transcriptional regulation in bacteria. We show that, in addition to these well-characterized protein families, the STAND class includes several other groups of (predicted) NTPase domains from diverse signaling and transcription regulatory proteins from bacteria and eukaryotes, and three Archaea-specific families. We identified the STAND domain in several biologically well-characterized proteins that have not been suspected to have NTPase activity, including soluble adenylyl cyclases, nephrocystin 3 (implicated in polycystic kidney disease), and Rolling pebble (a regulator of muscle development); these findings are expected to facilitate elucidation of the functions of these proteins. The STAND class belongs to the additional strand, catalytic E division of P-loop NTPases together with the AAA+ ATPases, RecA/helicase-related ATPases, ABC-ATPases, and VirD4/PilT-like ATPases. The STAND proteins are distinguished from other P-loop NTPases by the presence of unique sequence motifs associated with the N-terminal helix and the core strand-4, as well as a C-terminal helical bundle that is fused to the NTPase domain. This helical module contains a signature GxP motif in the loop between the two distal helices. With the exception of the archaeal families, almost all STAND NTPases are multidomain proteins containing three or more domains. In addition to the NTPase domain, these proteins typically contain DNA-binding or protein-binding domains, superstructure-forming repeats, such as WD40 and TPR, and enzymatic domains involved in signal transduction, including adenylate cyclases and kinases. By analogy to the AAA+ ATPases, it can be predicted that STAND NTPases use the C-terminal helical bundle as a "lever" to transmit the conformational changes brought about by NTP hydrolysis to effector domains. STAND NTPases represent a novel paradigm in signal transduction, whereby adaptor, regulatory switch, scaffolding, and, in some cases, signal-generating moieties are combined into a single polypeptide. The STAND class consists of 14 distinct families, and the evolutionary history of most of these families is riddled with dramatic instances of lineage-specific expansion and apparent horizontal gene transfer. The STAND NTPases are most abundant in developmentally and organizationally complex prokaryotes and eukaryotes. Transfer of genes for STAND NTPases from bacteria to eukaryotes on several occasions might have played a significant role in the evolution of eukaryotic signaling systems.

Adenosine Triphosphatases↗

Comparative genomics of the FtsK-HerA superfamily of pumping ATPases: implications for the origins of chromosome segregation, cell division and viral capsid packaging.

Recently, it has been shown that a predicted P-loop ATPase (the HerA or MlaA protein), which is highly conserved in archaea and also present in many bacteria but absent in eukaryotes, has a bidirectional helicase activity and forms hexameric rings similar to those described for the TrwB ATPase. In this study, the FtsK-HerA superfamily of P-loop ATPases, in which the HerA clade comprises one of the major branches, is analyzed in detail. We show that, in addition to the FtsK and HerA clades, this superfamily includes several families of characterized or predicted ATPases which are predominantly involved in extrusion of DNA and peptides through membrane pores. The DNA-packaging ATPases of various bacteriophages and eukaryotic double-stranded DNA viruses also belong to the FtsK-HerA superfamily. The FtsK protein is the essential bacterial ATPase that is responsible for the correct segregation of daughter chromosomes during cell division. The structural and evolutionary relationship between HerA and FtsK and the nearly perfect complementarity of their phyletic distributions suggest that HerA similarly mediates DNA pumping into the progeny cells during archaeal cell division. It appears likely that the HerA and FtsK families diverged concomitantly with the archaeal-bacterial division and that the last universal common ancestor of modern life forms had an ancestral DNA-pumping ATPase that gave rise to these families. Furthermore, the relationship of these cellular proteins with the packaging ATPases of diverse DNA viruses suggests that a common DNA pumping mechanism might be operational in both cellular and viral genome segregation. The herA gene forms a highly conserved operon with the gene for the NurA nuclease and, in many archaea, also with the orthologs of eukaryotic double-strand break repair proteins MRE11 and Rad50. HerA is predicted to function in a complex with these proteins in DNA pumping and repair of double-stranded breaks introduced during this process and, possibly, also during DNA replication. Extensive comparative analysis of the 'genomic context' combined with in-depth sequence analysis led to the prediction of numerous previously unnoticed nucleases of the NurA superfamily, including a specific version that is likely to be the endonuclease component of a novel restriction-modification system. This analysis also led to the identification of previously uncharacterized nucleases, such as a novel predicted nuclease of the Sir2-type Rossmann fold, and phosphatases of the HAD superfamily that are likely to function as partners of the FtsK-HerA superfamily ATPases.

Adenosine Triphosphatases↗

The SHS2 module is a common structural theme in functionally diverse protein groups, like Rpb7p, FtsA, GyrI, and MTH1598/TM1083 superfamilies.

Using structural comparisons, we identified a novel domain with a simple fold in the bacterial cell division ATPase FtsA, the archaeo-eukaryotic RNA polymerase subunit Rpb7p, the GyrI superfamily, and the uncharacterized MTH1598/Tm1083-like proteins. The fold contains a core of 3 strands, forming a curved sheet, and a single helix in a strand-helix-strand-strand (SHS2) configuration. The SHS2 domain may exist either in single or duplicate copies within the same polypeptide. The single-copy versions of the domain in FtsA and Rbp7p are most closely related, and appear to mediate protein-protein interactions by means of strand 1, and the loop between strand 2 and strand 3 of the domain. We predict that the interactions between FtsA and its functional partners in bacterial cell division are likely to be similar to the interactions of Rbp7p in the archaeo-eukaryotic RNA polymerase complex. The dimeric versions typified by the GyrI superfamily appear to have been adapted for small-molecule binding. Sequence profiles searches helped us to identify several new versions of the GyrI superfamily, including a family of secreted forms that is found only in animals and the bacterial pathogen Leptospira. Through sequence-structure comparisons, we predict the positions that are likely to be important for ligand specificity in the GyrI superfamily. In the MTH1598/Tm1083-like proteins, a SHS2 domain is inserted into the loop between strand 1 and helix 1 of another SHS2 domain. This has resulted in a structure that has convergent similarities with the Hsp33 and green fluorescent protein folds. The sequence conservation pattern and its phyletic profile suggest that it might function as an enzyme in some conserved aspect of nucleic acid metabolism. Thus, the SHS2 domain is an example of a simple module that has been adapted to perform an entire spectrum of functions ranging from protein-protein interactions to small-molecule recognition and catalysis.

Amino Acid Sequence↗

De-ubiquitination and ubiquitin ligase domains of A20 downregulate NF-kappaB signalling.

NF-kappaB transcription factors mediate the effects of pro-inflammatory cytokines such as tumour necrosis factor-alpha and interleukin-1beta. Failure to downregulate NF-kappaB transcriptional activity results in chronic inflammation and cell death, as observed in A20-deficient mice. A20 is a potent inhibitor of NF-kappaB signalling, but its mechanism of action is unknown. Here we show that A20 downregulates NF-kappaB signalling through the cooperative activity of its two ubiquitin-editing domains. The amino-terminal domain of A20, which is a de-ubiquitinating (DUB) enzyme of the OTU (ovarian tumour) family, removes lysine-63 (K63)-linked ubiquitin chains from receptor interacting protein (RIP), an essential mediator of the proximal TNF receptor 1 (TNFR1) signalling complex. The carboxy-terminal domain of A20, composed of seven C2/C2 zinc fingers, then functions as a ubiquitin ligase by polyubiquitinating RIP with K48-linked ubiquitin chains, thereby targeting RIP for proteasomal degradation. Here we define a novel ubiquitin ligase domain and identify two sequential mechanisms by which A20 downregulates NF-kappaB signalling. We also provide an example of a protein containing separate ubiquitin ligase and DUB domains, both of which participate in mediating a distinct regulatory effect.

Cell Line↗

Novel conserved domains in proteins with predicted roles in eukaryotic cell-cycle regulation, decapping and RNA stability.

BACKGROUND: The emergence of eukaryotes was characterized by the expansion and diversification of several ancient RNA-binding domains and the apparent de novo innovation of new RNA-binding domains. The identification of these RNA-binding domains may throw light on the emergence of eukaryote-specific systems of RNA metabolism. RESULTS: Using sensitive sequence profile searches, homology-based fold recognition and sequence-structure superpositions, we identified novel, divergent versions of the Sm domain in the Scd6p family of proteins. This family of Sm-related domains shares certain features of conventional Sm domains, which are required for binding RNA, in addition to possessing some unique conserved features. We also show that these proteins contain a second previously uncharacterized C-terminal domain, termed the FDF domain (after a conserved sequence motif in this domain). The FDF domain is also found in the fungal Dcp3p-like and the animal FLJ22128-like proteins, where it fused to a C-terminal domain of the YjeF-N domain family. In addition to the FDF domains, the FLJ22128-like proteins contain yet another divergent version of the Sm domain at their extreme N-terminus. We show that the YjeF-N domains represent a novel version of the Rossmann fold that has acquired a set of catalytic residues and structural features that distinguish them from the conventional dehydrogenases. CONCLUSIONS: Several lines of contextual information suggest that the Scd6p family and the Dcp3p-like proteins are conserved components of the eukaryotic RNA metabolism system. We propose that the novel domains reported here, namely the divergent versions of the Sm domain and the FDF domain may mediate specific RNA-protein and protein-protein interactions in cytoplasmic ribonucleoprotein complexes. More specifically, the protein complexes containing Sm-like domains of the Scd6p family are predicted to regulate the stability of mRNA encoding proteins involved in cell cycle progression and vesicular assembly. The Dcp3p and FLJ22128 proteins may localize to the cytoplasmic processing bodies and possibly catalyze a specific processing step in the decapping pathway. The explosive diversification of Sm domains appears to have played a role in the emergence of several uniquely eukaryotic ribonucleoprotein complexes, including those involved in decapping and mRNA stability.

Amino Acid Sequence↗

Evolution of bacterial RNA polymerase: implications for large-scale bacterial phylogeny, domain accretion, and horizontal gene transfer.

Comparative analysis of the domain architectures of the beta, beta', and sigma(70) subunits of bacterial DNA-dependent RNA polymerases (DdRp), combined with sequence-based phylogenetic analysis, revealed a fundamental split among bacteria. DNA-dependent RNA polymerase subunits of Group I, which includes Proteobacteria, Aquifex, Chlamydia, Spirochaetes, Cytophaga-Chlorobium, and Planctomycetes, are characterized by three distinct inserts, namely a Sandwich Barrel Hybrid Motif domain in the beta subunit, a beta-beta' module (BBM) 1 domain in the beta' subunit, and a distinct helical module in the sigma subunit. The DdRp subunits of remaining bacteria, which comprise Group II, lack these inserts, although some additional inserted domains are present in individual lineages. The separation of bacteria into Group I and Group II is generally compatible with the topologies of phylogenetic trees of the conserved regions of DdRp subunits and concatenated ribosomal proteins and might represent the primary bifurcation in bacterial evolution. A striking deviation from this evolutionary pattern is Aquifex whose DdRp subunits cluster within Group I, whereas phylogenetic analysis of ribosomal proteins identifies Aquifex as grouping with Thermotoga another bacterial hyperthemophile belonging to Group II. The inferred evolutionary scenario for the DdRp subunits includes domain accretion and rearrangement, with some likely horizontal transfer events. Although evolution of bacterial DdRp appeared to be generally dominated by vertical inheritance, horizontal transfer of complete genes for all or some of the subunits, resulting in displacement of the ancestral genes, might have played a role in several lineages, such as Aquifex, Thermotoga, and Fusobacterium.

Amino Acid Sequence↗

The C. elegans Polycomb gene SOP-2 encodes an RNA binding protein.

Epigenetic silencing of Hox cluster genes by Polycomb group (PcG) proteins is thought to involve the formation of a stably inherited repressive chromatin structure. Here we show that the C. elegans-specific PcG protein SOP-2 directly binds to RNA through three nonoverlapping regions, each of which is essential for its localization to characteristic nuclear bodies and for its in vivo function in the repression of Hox genes. Functional studies indicate that the RNA involved in SOP-2 binding is distinct from either siRNA or microRNA. Remarkably, the vertebrate PcG protein Rae28, which is functionally and structurally related to SOP-2, also binds to RNA through an FCS finger domain. Substitution of the Rae28 FCS finger for the essential RNA binding region of SOP-2 partially restores localization to nuclear bodies. These observations suggest that direct binding to RNA is an evolutionarily conserved and potentially important property of PcG proteins.

Animals↗

A multidomain adhesion protein family expressed in Plasmodium falciparum is essential for transmission to the mosquito.

The recent sequencing of several apicomplexan genomes has provided the opportunity to characterize novel antigens essential for the parasite life cycle that might lead to the development of new diagnostic and therapeutic markers. Here we have screened the Plasmodium falciparum genome sequence for genes encoding extracellular multidomain putative adhesive proteins. Three of these identified genes, named PfCCp1, PfCCp2, and PfCCp3, have multiple adhesive modules including a common Limulus coagulation factor C domain also found in two additional Plasmodium genes. Orthologues were identified in the Cryptosporidium parvum genome sequence, indicating an evolutionary conserved function. Transcript and protein expression analysis shows sexual stage-specific expression of PfCCp1, PfCCp2, and PfCCp3, and cellular localization studies revealed plasma membrane-associated expression in mature gametocytes. During gametogenesis, PfCCps are released and localize surrounding complexes of newly emerged microgametes and macrogametes. PfCCp expression markedly decreased after formation of zygotes. To begin to address PfCCp function, the PfCCp2 and PfCCp3 gene loci were disrupted by homologous recombination, resulting in parasites capable of forming oocyst sporozoites but blocked in the salivary gland transition. Our results describe members of a conserved apicomplexan protein family expressed in sexual stage Plasmodium parasites that may represent candidates for subunits of a transmission-blocking vaccine.

Amino Acid Sequence↗

The emergence of catalytic and structural diversity within the beta-clip fold.

The beta-clip fold includes a diverse group of protein domains that are unified by the presence of two characteristic waist-like constrictions, which bound a central extended region. Members of this fold include enzymes like deoxyuridine triphosphatase and the SET methylase, carbohydrate-binding domains like the fish antifreeze proteins/Sialate synthase C-terminal domains, and functionally enigmatic accessory subunits of urease and molybdopterin biosynthesis protein MoeA. In this study, we reconstruct the evolutionary history of this fold using sensitive sequence and structure comparisons methods. Using sequence profile searches, we identified novel versions of the beta-clip fold in the bacterial flagellar chaperone FlgA and the related pilus protein CpaB, the StrU-like dehydrogenases, and the UxaA/GarD-like hexuronate dehydratases (SAF superfamily). We present evidence that these versions of the beta-clip domain, like the related type III anti-freeze proteins and C-terminal domains of sialic acid synthases, are involved in interactions with carbohydrates. We propose that the FlgA and CpaB-like proteins mediate the assembly of bacterial flagella and Flp pili by means of their interactions with the carbohydrate moieties of peptidoglycan. The N-terminal beta-clip domain of the hexuronate dehydratases appears to have evolved a novel metal-binding site, while their C-terminal domain is likely to adopt a metal-binding TIM barrel-like fold. Using structural comparisons, we show that the beta-clip fold can be further classified into two major groups, one that includes the SAF, SET, dUTPase superfamilies, and the other that includes the phage lambda head decoration protein, the beta subunit of urease and the C-terminal domain of the molybdenum cofactor biosynthesis protein MoeA. Structural comparisons also suggest the beta-clip fold was assembled through the duplication of a three-stranded unit. Though the three-stranded units are likely to have had a common origin, we present evidence that complete beta-clip domains were assembled through such duplications, independently on multiple occasions. There is also evidence for circular permutation of the basic three-stranded unit on different occasions in the evolution of the beta-clip unit. We also describe how assembly of this fold from a basic three-stranded unit has been utilized to accommodate a variety of activities in its different versions.

Amino Acid Sequence↗