Search PubMed⌕ Search

Biomedical subjects

L Aravind

Publications and source records attributed to L Aravind.

At least 73 records · Page 4Linked to original sources

Evolutionary connections between bacterial and eukaryotic signaling systems: a genomic perspective.

Recent advances in microbial genomics suggest that several protein domains are common to bacterial and eukaryotic regulatory proteins. In particular, developmentally and morphologically complex prokaryotes appear to share several signaling modules with eukaryotes. New experimental studies and information from domain architectures point to several similar mechanistic themes in bacterial and eukaryotic signaling proteins. Laterally transferred protein domains, originally of bacterial provenance, appear to have contributed to the evolution of sensory pathways related to light, redox and nitric oxide signaling, and developmental pathways, such as Notch, cytokine and cytokinin signaling in eukaryotes.

Bacterial Physiological Phenomena↗

Emergence of diverse biochemical activities in evolutionarily conserved structural scaffolds of proteins.

Comparative analysis of numerous protein structures that have become available in the past few years, combined with genome comparison, has yielded new insights into the evolution of enzymes and their functions. In addition to the well-known diversification of substrate specificities, enzymes with several widespread catalytic folds, particularly the TIM barrel, the RRM-like domain and the double-stranded beta-helix (cupin) domain, have been extensively explored in 'reaction space', resulting in the evolution of numerous, diverse catalytic activities supported by the same structural scaffold. Common protein folds differ widely in the diversity of catalyzed reactions. The biochemical plasticity of a fold seems to hinge on the presence of a generic, symmetrical substrate-binding pocket as opposed to highly specialized binding sites.

Catalysis↗

Multiple transporters associated with malaria parasite responses to chloroquine and quinine.

Mutations and/or overexpression of various transporters are known to confer drug resistance in a variety of organisms. In the malaria parasite Plasmodium falciparum, a homologue of P-glycoprotein, PfMDR1, has been implicated in responses to chloroquine (CQ), quinine (QN) and other drugs, and a putative transporter, PfCRT, was recently demonstrated to be the key molecule in CQ resistance. However, other unknown molecules are probably involved, as different parasite clones carrying the same pfcrt and pfmdr1 alleles show a wide range of quantitative responses to CQ and QN. Such molecules may contribute to increasing incidences of QN treatment failure, the molecular basis of which is not understood. To identify additional genes involved in parasite CQ and QN responses, we assayed the in vitro susceptibilities of 97 culture-adapted cloned isolates to CQ and QN and searched for single nucleotide polymorphisms (SNPs) in DNA encoding 49 putative transporters (total 113 kb) and in 39 housekeeping genes that acted as negative controls. SNPs in 11 of the putative transporter genes, including pfcrt and pfmdr1, showed significant associations with decreased sensitivity to CQ and/or QN in P. falciparum. Significant linkage disequilibria within and between these genes were also detected, suggesting interactions among the transporter genes. This study provides specific leads for better understanding of complex drug resistances in malaria parasites.

Animals↗

Gene duplication with displacement and rearrangement: origin of the bacterial replication protein PriB from the single-stranded DNA-binding protein Ssb.

PriB is a proteobacterial protein that is involved in the pre-primosomal step of DNA replication and, unexpectedly, is encoded with a ribosomal protein operon. Detailed sequence comparisons and analysis of operon organization show that PriB evolved from the single-stranded DNA-binding (Ssb) via gene duplication with subsequent rapid sequence diversification. Duplication of the SSB gene was accompanied by a genome rearrangement which resulted in one of the paralogs retaining the original position, whereas the other was relocated. The functional specialization of the resulting paralogs apparently proceeded in an unexpected fashion: the original Ssb function remained with the relocated paralog, whereas the one within the ribosomal protein operon acquired a new, specialized function in replication.

Amino Acid Sequence↗

Detection of novel members, structure-function analysis and evolutionary classification of the 2H phosphoesterase superfamily.

2',3' Cyclic nucleotide phosphodiesterases are enzymes that catalyze at least two distinct steps in the splicing of tRNA introns in eukaryotes. Recently, the biochemistry and structure of these enzymes, from yeast and the plant Arabidopsis thaliana, have been extensively studied. They were found to share a common active site, characterized by two conserved histidines, with the bacterial tRNA-ligating enzyme LigT and the vertebrate myelin-associated 2',3' phosphodiesterases. Using sensitive sequence profile analysis methods, we show that these enzymes define a large superfamily of predicted phosphoesterases with two conserved histidines (hence 2H phosphoesterase superfamily). We identify several new families of 2H phosphoesterases and present a complete evolutionary classification of this superfamily. We also carry out a structure- function analysis of these proteins and present evidence for diverse interactions for different families, within this superfamily, with RNA substrates and protein partners. In particular, we show that eukaryotes contain two ancient families of these proteins that might be involved in RNA processing, transcriptional co-activation and post-transcriptional gene silencing. Another eukaryotic family restricted to vertebrates and insects is combined with UBA and SH3 domains suggesting a role in signal transduction. We detect these phosphoesterase modules in polyproteins of certain retroviruses, rotaviruses and coronaviruses, where they could function in capping and processing of viral RNAs. Furthermore, we present evidence for multiple families of 2H phosphoesterases in bacteria, which might be involved in the processing of small molecules with the 2',3' cyclic phosphoester linkages. The evolutionary analysis suggests that the 2H domain emerged through a duplication of a simple structural unit containing a single catalytic histidine prior to the last common ancestor of all life forms. Initially, this domain appears to have been involved in RNA processing and it appears to have been recruited to perform various other functions in later stages of evolution.

2',3'-Cyclic-Nucleotide Phosphodiesterases↗

The catalytic domains of thiamine triphosphatase and CyaB-like adenylyl cyclase define a novel superfamily of domains that bind organic phosphates.

BACKGROUND: The CyaB protein from Aeromonas hydrophila has been shown to possess adenylyl cyclase activity. While orthologs of this enzyme have been found in some bacteria and archaea, it shows no detectable relationship to the classical nucleotide cyclases. Furthermore, the actual biological functions of these proteins are not clearly understood because they are also present in organisms in which there is no evidence for cyclic nucleotide signaling. RESULTS: We show that the CyaB like adenylyl cyclase and the mammalian thiamine triphosphatases define a novel superfamily of catalytic domains called the CYTH domain that is present in all three superkingdoms of life. Using multiple alignments and secondary structure predictions, we define the catalytic core of these enzymes to contain a novel alpha+beta scaffold with 6 conserved acidic residues and 4 basic residues. Using contextual information obtained from the analysis of gene neighborhoods and domain fusions, we predict that members of this superfamily may play a central role in the interface between nucleotide and polyphosphate metabolism. Additionally, based on contextual information, we identify a novel domain (called CHAD) that is predicted to functionally interact with the CYTH domain-containing enzymes in bacteria and archaea. The CHAD is predicted to be an alpha helical domain, and contains conserved histidines that may be critical for its function. CONCLUSIONS: The phyletic distribution of the CYTH domain suggests that it is an ancient enzymatic domain that was present in the Last Universal Common Ancestor and was involved in nucleotide or organic phosphate metabolism. Based on the conservation of catalytic residues, we predict that CYTH domains are likely to chelate two divalent cations, and exhibit a reaction mechanism that is dependent on two metal ions, analogous to nucleotide cyclases, polymerases and certain phosphoesterases. Our analysis also suggests that the experimentally characterized members of this superfamily, namely adenylyl cyclase and thiamine triphosphatase, are secondary derivatives of proteins that performed an ancient role in polyphosphate and nucleotide metabolism.

Journal Article↗

The PRC-barrel: a widespread, conserved domain shared by photosynthetic reaction center subunits and proteins of RNA metabolism.

BACKGROUND: The H subunit of the purple bacterial photosynthetic reaction center (PRC-H) is important for the assembly of the photosynthetic reaction center and appears to regulate electron transfer during the reduction of the secondary quinone. It contains a distinct cytoplasmic beta-barrel domain whose fold has no close structural relationship to any other well known beta-barrel domain. RESULTS: We show that the PRC-H beta-barrel domain is the prototype of a novel superfamily of protein domains, the PRC-barrels, approximately 80 residues long, which is widely represented in bacteria, archaea and plants. This domain is also present at the carboxyl terminus of the pan-bacterial protein RimM, which is involved in ribosomal maturation and processing of 16S rRNA. A family of small proteins conserved in all known euryarchaea are composed entirely of a single stand-alone copy of the domain. Versions of this domain from photosynthetic proteobacteria contain a conserved acidic residue that is thought to regulate the reduction of quinones in the light-induced electron-transfer reaction. Closely related forms containing this acidic residue are also found in several non-photosynthetic bacteria, as well as in cyanobacteria, which have reaction centers with a different organization. We also show that the domain contains several determinants that could mediate specific protein-protein interactions. CONCLUSIONS: The PRC-barrel is a widespread, ancient domain that appears to have been recruited to a variety of biological systems, ranging from RNA processing to photosynthesis. Identification of this versatile domain in numerous proteins could aid investigation of unexplored aspects of their biology.

Amino Acid Sequence↗

YjeQ, an essential, conserved, uncharacterized protein from Escherichia coli, is an unusual GTPase with circularly permuted G-motifs and marked burst kinetics.

The Escherichia coli protein YjeQ represents a protein family whose members are broadly conserved in bacteria and have been shown to be indispensable to the growth of E. coli and Bacillus subtilis [Arigoni, F., et al. (1998) Nat. Biotechnol. 16, 851]. Proteins of the YjeQ family contain all sequence motifs typical of the vast class of P-loop-containing GTPases, but show a circular permutation, with a G4-G1-G3 pattern of motifs as opposed to the regular G1-G3-G4 pattern seen in most GTPases. All YjeQ family proteins display a unique domain architecture, which includes a predicted N-terminal OB-fold RNA-binding domain, the central permuted GTPase module, and a zinc knuckle-like C-terminal cysteine cluster. This domain architecture suggests a possible role for YjeQ as a regulator of translation. YjeQ was overexpressed, purified to homogeneity, and shown to contain 0.6 equiv of GDP. Steady state kinetic analyses indicated slow GTP hydrolysis, with a k(cat) of 9.4 h(-)(1) and a K(m) for GTP of 120 microM (k(cat)/K(m) = 21.7 M(-)(1) s(-)(1)). YjeQ also hydrolyzed other nucleoside triphosphates and deoxynucleotide triphosphates such as ATP, ITP, and CTP with specificity constants (k(cat)/K(m)) ranging from 0.2 to 1.0 M(-)(1) s(-)(1). Pre-steady state kinetic analysis of YjeQ revealed a burst of nucleotide hydrolysis for GTP described by a first-order rate constant of 100 s(-)(1) as compared to a burst rate of 0.2 s(-)(1) for ATP. In addition, a variant in the G1 motif of YjeQ (S221A) was substantially impaired for GTP hydrolysis (0.3 s(-)(1)) with a less significant impact on the steady state rate (1.8 h(-)(1)). In summary, E. coli YjeQ is an unusual, circularly permuted P-loop-containing GTPase, which catalyzes GTP hydrolysis at a rate 45 000 times greater than that of turnover.

Amino Acid Motifs↗

Role of Rpn11 metalloprotease in deubiquitination and degradation by the 26S proteasome.

The 26S proteasome mediates degradation of ubiquitin-conjugated proteins. Although ubiquitin is recycled from proteasome substrates, the molecular basis of deubiquitination at the proteasome and its relation to substrate degradation remain unknown. The Rpn11 subunit of the proteasome lid subcomplex contains a highly conserved Jab1/MPN domain-associated metalloisopeptidase (JAMM) motif-EX(n)HXHX(10)D. Mutation of the predicted active-site histidines to alanine (rpn11AXA) was lethal and stabilized ubiquitin pathway substrates in yeast. Rpn11(AXA) mutant proteasomes assembled normally but failed to either deubiquitinate or degrade ubiquitinated Sic1 in vitro. Our findings reveal an unexpected coupling between substrate deubiquitination and degradation and suggest a unifying rationale for the presence of the lid in eukaryotic proteasomes.

Adenosine Triphosphate↗

Role of predicted metalloprotease motif of Jab1/Csn5 in cleavage of Nedd8 from Cul1.

COP9 signalosome (CSN) cleaves the ubiquitin-like protein Nedd8 from the Cul1 subunit of SCF ubiquitin ligases. The Jab1/MPN domain metalloenzyme (JAMM) motif in the Jab1/Csn5 subunit was found to underlie CSN's Nedd8 isopeptidase activity. JAMM is found in proteins from archaea, bacteria, and eukaryotes, including the Rpn11 subunit of the 26S proteasome. Metal chelators and point mutations within JAMM abolished CSN-dependent cleavage of Nedd8 from Cul1, yet had little effect on CSN complex assembly. Optimal SCF activity in yeast and both viability and proper photoreceptor cell (R cell) development in Drosophila melanogaster required an intact Csn5 JAMM domain. We propose that JAMM isopeptidases play important roles in a variety of physiological pathways.

Amino Acid Motifs↗

The SWIRM domain: a conserved module found in chromosomal proteins points to novel chromatin-modifying activities.

BACKGROUND: Eukaryotic chromosomal components, especially histones, are subject to a wide array of covalent modifications and catalytic reorganization. These modifications have an important role in the regulation of chromatin structure and are mediated by large multisubunit complexes that contain modular proteins with several conserved catalytic and noncatalytic adaptor domains. RESULTS: Using computational sequence-profile analysis methods, we identified a previously uncharacterized, predicted alpha-helical domain of about 85 residues in chromosomal proteins such as Swi3p, Rsc8p, Moira and several other uncharacterized proteins. This module, termed the SWIRM domain, is predicted to mediate specific protein-protein interactions in the assembly of chromatin-protein complexes. In one group of proteins, which are highly conserved throughout the crown-group eukaryotes, the SWIRM domain is linked to a catalytic domain related to the monoamine and polyamine oxidases. Another human protein has the SWIRM domain linked to a JAB domain that is involved in protein degradation through the ubiquitin pathway. CONCLUSIONS: Identification of the SWIRM domain could help in directed experimental analysis of specific interactions in chromosomal proteins. We predict that the proteins in which it is combined with an amino-oxidase domain define a novel class of chromatin-modifying enzymes, which are likely to oxidize either the amino group of basic residues in histones and other chromosomal proteins or the polyamines in chromatin, and thereby alter the charge distribution. Other forms, such as KIAA1915, may link chromatin modification to ubiquitin-dependent protein degradation.

Amino Acid Motifs↗

Monophyly of class I aminoacyl tRNA synthetase, USPA, ETFP, photolyase, and PP-ATPase nucleotide-binding domains: implications for protein evolution in the RNA.

Protein sequence and structure comparisons show that the catalytic domains of Class I aminoacyl-tRNA synthetases, a related family of nucleotidyltransferases involved primarily in coenzyme biosynthesis, nucleotide-binding domains related to the UspA protein (USPA domains), photolyases, electron transport flavoproteins, and PP-loop-containing ATPases together comprise a distinct class of alpha/beta domains designated the HUP domain after HIGH-signature proteins, UspA, and PP-ATPase. Several lines of evidence are presented to support the monophyly of the HUP domains, to the exclusion of other three-layered alpha/beta folds with the generic "Rossmann-like" topology. Cladistic analysis, with patterns of structural and sequence similarity used as discrete characters, identified three major evolutionary lineages within the HUP domain class: the PP-ATPases; the HIGH superfamily, which includes class I aaRS and related nucleotidyltransferases containing the HIGH signature in their nucleotide-binding loop; and a previously unrecognized USPA-like group, which includes USPA domains, electron transport flavoproteins, and photolyases. Examination of the patterns of phyletic distribution of distinct families within these three major lineages suggests that the Last Universal Common Ancestor of all modern life forms encoded 15-18 distinct alpha/beta ATPases and nucleotide-binding proteins of the HUP class. This points to an extensive radiation of HUP domains before the last universal common ancestor (LUCA), during which the multiple class I aminoacyl-tRNA synthetases emerged only at a late stage. Thus, substantial evolutionary diversification of protein domains occurred well before the modern version of the protein-dependent translation machinery was established, i.e., still in the RNA world.

Adenosine Triphosphatases↗

The GOLD domain, a novel protein module involved in Golgi function and secretion.

BACKGROUND: Members of the p24 (p24/gp25L/emp24/Erp) family of proteins have been shown to be critical components of the coated vesicles that are involved in the transportation of cargo molecules from the endoplasmic reticulum to the Golgi complex. The p24 proteins form hetero-oligomeric complexes and are believed to function as receptors for specific secretory cargo. RESULTS: Using sensitive sequence-profile analysis methods, we identified a novel beta-strand-rich domain, the GOLD (Golgi dynamics) domain, in the p24 proteins and several other proteins with roles in Golgi dynamics and secretion. This domain is predicted to mediate diverse protein-protein interactions. Other than in the p24 proteins, the GOLD domain is always found combined with lipid- or membrane-association domains such as the pleckstrin homology (PH), Sec14p and FYVE domains. CONCLUSIONS: The identification of the GOLD domain could aid in directed investigation of the role of the p24 proteins in the secretion process. The newly detected group of GOLD-domain proteins, which might simultaneously bind membranes and other proteins, point to the existence of a novel class of adaptors that could have a role in the assembly of membrane-associated complexes or in regulating assembly of cargo into membranous vesicles.

Amino Acid Sequence↗

The complete genome of hyperthermophile Methanopyrus kandleri AV19 and monophyly of archaeal methanogens.

We have determined the complete 1,694,969-nt sequence of the GC-rich genome of Methanopyrus kandleri by using a whole direct genome sequencing approach. This approach is based on unlinking of genomic DNA with the ThermoFidelase version of M. kandleri topoisomerase V and cycle sequencing directed by 2'-modified oligonucleotides (Fimers). Sequencing redundancy (3.3x) was sufficient to assemble the genome with less than one error per 40 kb. Using a combination of sequence database searches and coding potential prediction, 1,692 protein-coding genes and 39 genes for structural RNAs were identified. M. kandleri proteins show an unusually high content of negatively charged amino acids, which might be an adaptation to the high intracellular salinity. Previous phylogenetic analysis of 16S RNA suggested that M. kandleri belonged to a very deep branch, close to the root of the archaeal tree. However, genome comparisons indicate that, in both trees constructed using concatenated alignments of ribosomal proteins and trees based on gene content, M. kandleri consistently groups with other archaeal methanogens. M. kandleri shares the set of genes implicated in methanogenesis and, in part, its operon organization with Methanococcus jannaschii and Methanothermobacter thermoautotrophicum. These findings indicate that archaeal methanogens are monophyletic. A distinctive feature of M. kandleri is the paucity of proteins involved in signaling and regulation of gene expression. Also, M. kandleri appears to have fewer genes acquired via lateral transfer than other archaea. These features might reflect the extreme habitat of this organism.

Base Sequence↗

Comparative genomics and evolution of proteins involved in RNA metabolism.

RNA metabolism, broadly defined as the compendium of all processes that involve RNA, including transcription, processing and modification of transcripts, translation, RNA degradation and its regulation, is the central and most evolutionarily conserved part of cell physiology. A comprehensive, genome-wide census of all enzymatic and non-enzymatic protein domains involved in RNA metabolism was conducted by using sequence profile analysis and structural comparisons. Proteins related to RNA metabolism comprise from 3 to 11% of the complete protein repertoire in bacteria, archaea and eukaryotes, with the greatest fraction seen in parasitic bacteria with small genomes. Approximately one-half of protein domains involved in RNA metabolism are present in most, if not all, species from all three primary kingdoms and are traceable to the last universal common ancestor (LUCA). The principal features of LUCA's RNA metabolism system were reconstructed by parsimony-based evolutionary analysis of all relevant groups of orthologous proteins. This reconstruction shows that LUCA possessed not only the basal translation system, but also the principal forms of RNA modification, such as methylation, pseudouridylation and thiouridylation, as well as simple mechanisms for polyadenylation and RNA degradation. Some of these ancient domains form paralogous groups whose evolution can be traced back in time beyond LUCA, towards low-specificity proteins, which probably functioned as cofactors for ribozymes within the RNA world framework. The main lineage-specific innovations of RNA metabolism systems were identified. The most notable phase of innovation in RNA metabolism coincides with the advent of eukaryotes and was brought about by the merge of the archaeal and bacterial systems via mitochondrial endosymbiosis, but also involved emergence of several new, eukaryote-specific RNA-binding domains. Subsequent, vast expansions of these domains mark the origin of alternative splicing in animals and probably in plants. In addition to the reconstruction of the evolutionary history of RNA metabolism, this analysis produced numerous functional predictions, e.g. of previously undetected enzymes of RNA modification.

Animals↗

Classification and evolutionary history of the single-strand annealing proteins, RecT, Redbeta, ERF and RAD52.

BACKGROUND: The DNA single-strand annealing proteins (SSAPs), such as RecT, Redbeta, ERF and Rad52, function in RecA-dependent and RecA-independent DNA recombination pathways. Recently, they have been shown to form similar helical quaternary superstructures. However, despite the functional similarities between these diverse SSAPs, their actual evolutionary affinities are poorly understood. RESULTS: Using sensitive computational sequence analysis, we show that the RecT and Redbeta proteins, along with several other bacterial proteins, form a distinct superfamily. The ERF and Rad52 families show no direct evolutionary relationship to these proteins and define novel superfamilies of their own. We identify several previously unknown members of each of these superfamilies and also report, for the first time, bacterial and viral homologs of Rad52. Additionally, we predict the presence of aberrant HhH modules in RAD52 that are likely to be involved in DNA-binding. Using the contextual information obtained from the analysis of gene neighborhoods, we provide evidence of the interaction of the bacterial members of each of these SSAP superfamilies with a similar set of DNA repair/recombination protein. These include different nucleases or Holliday junction resolvases, the ABC ATPase SbcC and the single-strand-binding protein. We also present evidence of independent assembly of some of the predicted operons encoding SSAPs and in situ displacement of functionally similar genes. CONCLUSIONS: There are three evolutionarily distinct superfamilies of SSAPs, namely the RecT/Redbeta, ERF, and RAD52, that have different sequence conservation patterns and predicted folds. All these SSAPs appear to be primarily of bacteriophage origin and have been acquired by numerous phylogenetically distant cellular genomes. They generally occur in predicted operons encoding one or more of a set of conserved DNA recombination proteins that appear to be the principal functional partners of the SSAPs.

Journal Article↗

Classification and evolution of P-loop GTPases and related ATPases.

Sequences and available structures were compared for all the widely distributed representatives of the P-loop GTPases and GTPase-related proteins with the aim of constructing an evolutionary classification for this superclass of proteins and reconstructing the principal events in their evolution. The GTPase superclass can be divided into two large classes, each of which has a unique set of sequence and structural signatures (synapomorphies). The first class, designated TRAFAC (after translation factors) includes enzymes involved in translation (initiation, elongation, and release factors), signal transduction (in particular, the extended Ras-like family), cell motility, and intracellular transport. The second class, designated SIMIBI (after signal recognition particle, MinD, and BioD), consists of signal recognition particle (SRP) GTPases, the assemblage of MinD-like ATPases, which are involved in protein localization, chromosome partitioning, and membrane transport, and a group of metabolic enzymes with kinase or related phosphate transferase activity. These two classes together contain over 20 distinct families that are further subdivided into 57 subfamilies (ancient lineages) on the basis of conserved sequence motifs, shared structural features, and domain architectures. Ten subfamilies show a universal phyletic distribution compatible with presence in the last universal common ancestor of the extant life forms (LUCA). These include four translation factors, two OBG-like GTPases, the YawG/YlqF-like GTPases (these two subfamilies also consist of predicted translation factors), the two signal-recognition-associated GTPases, and the MRP subfamily of MinD-like ATPases. The distribution of nucleotide specificity among the proteins of the GTPase superclass indicates that the common ancestor of the entire superclass was a GTPase and that a secondary switch to ATPase activity has occurred on several independent occasions during evolution. The functions of most GTPases that are traceable to LUCA are associated with translation. However, in contrast to other superclasses of P-loop NTPases (RecA-F1/F0, AAA+, helicases, ABC), GTPases do not participate in NTP-dependent nucleic acid unwinding and reorganizing activities. Hence, we hypothesize that the ancestral GTPase was an enzyme with a generic regulatory role in translation, with subsequent diversification resulting in acquisition of diverse functions in transport, protein trafficking, and signaling. In addition to the classification of previously known families of GTPases and related ATPases, we introduce several previously undetected families and describe new functional predictions.

Adenosine Triphosphatases↗

Classification of the caspase-hemoglobinase fold: detection of new families and implications for the origin of the eukaryotic separins.

A comprehensive sequence and structural comparative analysis of the caspase-hemoglobinase protein fold resulted in the delineation of the minimal structural core of the protease domain and the identification of numerous, previously undetected members, including a new protease family typified by the HetF protein from the cyanobacterium Nostoc. The first bacterial homologs of legumains and hemoglobinases were also identified. Most proteins containing this fold are known or predicted to be active proteases, but multiple, independent inactivations were noticed in nearly all lineages. Together with the tendency of caspase-related proteases to form intramolecular or intermolecular dimers, this suggests a widespread regulatory role for the inactive forms. A classification of the caspase-hemoglobinase fold was developed to reflect the inferred evolutionary relationships between the constituent protein families. Proteins containing this domain were so far detected almost exclusively in bacteria and eukaryotes. This analysis indicates that caspase-hemoglobinase-fold proteases and their inactivated derivatives are widespread in diverse bacteria, particularly those with a complex development, such as Streptomyces, Anabaena, Mesorhizobium, and Myxococcus. The eukaryotic separin family was shown to be most closely related to the mainly prokaryotic HetF family. The phyletic patterns and evolutionary relationships between these proteins suggest that they probably were acquired by eukaryotes from bacteria during the primary, promitochondrial endosymbiosis. A similar scenario, supported by phylogenetic analysis, seems to apply to metacaspases and paracaspases, with the latter, perhaps, being acquired in an independent horizontal transfer to the eukaryotes. The acquisition of the caspase-hemoglobinase-fold domains by eukaryotes might have been critical in the evolution of important eukaryotic processes, such as mitosis and programmed cell death.

Amino Acid Sequence↗