Search PubMed⌕ Search

Biomedical subjects

L Aravind

Publications and source records attributed to L Aravind.

At least 145 records · Page 8Linked to original sources

Eukaryote-specific domains in translation initiation factors: implications for translation regulation and evolution of the translation system.

Computational analysis of sequences of proteins involved in translation initiation in eukaryotes reveals a number of specific domains that are not represented in bacteria or archaea. Most of these eukaryote-specific domains are known or predicted to possess an alpha-helical structure, which suggests that such domains are easier to invent in the course of evolution than are domains of other structural classes. A previously undetected, conserved region predicted to form an alpha-helical domain is delineated in the initiation factor eIF4G, in Nonsense-mediated mRNA decay 2 protein (NMD2/UPF2), in the nuclear cap-binding CBP80, and in other, poorly characterized proteins, which is named the NIC (NMD2, eIF4G, CBP80) domain. Biochemical and mutagenesis data on NIC-containing proteins indicate that this predicted domain is one of the central adapters in the regulation of mRNA processing, translation, and degradation. It is demonstrated that, in the course of eukaryotic evolution, initiation factor eIF4G, of which NIC is the core, conserved portion, has accreted several additional, distinct predicted domains such as MI (MA-3 and eIF4G ) and W2, which probably was accompanied by acquisition of new regulatory interactions.

Adaptor Proteins, Signal Transducing↗

Comparative genome analysis of the pathogenic spirochetes Borrelia burgdorferi and Treponema pallidum.

A comparative analysis of the predicted protein sequences encoded in the complete genomes of Borrelia burgdorferi and Treponema pallidum provides a number of insights into evolutionary trends and adaptive strategies of the two spirochetes. A measure of orthologous relationships between gene sets, termed the orthology coefficient (OC), was developed. The overall OC value for the gene sets of the two spirochetes is about 0.43, which means that less than one-half of the genes show readily detectable orthologous relationships. This emphasizes significant divergence between the two spirochetes, apparently driven by different biological niches. Different functional categories of proteins as well as different protein families show a broad distribution of OC values, from near 1 (a perfect, one-to-one correspondence) to near 0. The proteins involved in core biological functions, such as genome replication and expression, typically show high OC values. In contrast, marked variability is seen among proteins that are involved in specific processes, such as nutrient transport, metabolism, gene-specific transcription regulation, signal transduction, and host response. Differences in the gene complements encoded in the two spirochete genomes suggest active adaptive evolution for their distinct niches. Comparative analysis of the spirochete genomes produced evidence of gene exchanges with other bacteria, archaea, and eukaryotic hosts that seem to have occurred at different points in the evolution of the spirochetes. Examples are presented of the use of sequence profile analysis to predict proteins that are likely to play a role in pathogenesis, including secreted proteins that contain specific protein-protein interaction domains, such as von Willebrand A, YWTD, TPR, and PR1, some of which hitherto have been reported only in eukaryotes. We tentatively reconstruct the likely evolutionary process that has led to the divergence of the two spirochete lineages; this reconstruction seems to point to an ancestral state resembling the symbiotic spirochetes found in insect guts.

Amino Acid Sequence↗

The bacterial replicative helicase DnaB evolved from a RecA duplication.

The RecA/Rad51/DCM1 family of ATP-dependent recombinases plays a crucial role in genetic recombination and double-stranded DNA break repair in Archaea, Bacteria, and Eukaryota. DnaB is the replication fork helicase in all Bacteria. We show here that DnaB shares significant sequence similarity with RecA and Rad51/DMC1 and two other related families of ATPases, Sms and KaiC. The conserved region spans the entire ATP- and DNA-binding domain that consists of about 250 amino acid residues and includes 7 distinct motifs. Comparison with the three-dimensional structure of Escherichia coli RecA and phage T7 DnaB (gp4) reveals that the area of sequence conservation includes the central parallel beta-sheet and most of the connecting helices and loops as well as a smaller domain that consists of a amino-terminal helix and a carboxy-terminal beta-meander. Additionally, we show that animals, plants, and the malarial Plasmodium but not Saccharomyces cerevisiae encode a previously undetected DnaB homolog that might function in the mitochondria. The DnaB homolog from Arabidopsis also contains a DnaG-primase domain and the DnaB homolog from the nematode seems to contain an inactivated version of the primase. This domain organization is reminiscent of bacteriophage primases-helicases and suggests that DnaB might have been horizontally introduced into the nuclear eukaryotic genome via a phage vector. We hypothesize that DnaB originated from a duplication of a RecA-like ancestor after the divergence of the bacteria from Archaea and eukaryotes, which indicates that the replication fork helicases in Bacteria and Archaea/Eukaryota have evolved independently.

Amino Acid Sequence↗

DNA-binding proteins and evolution of transcription regulation in the archaea.

Likely DNA-binding domains in archaeal proteins were analyzed using sequence profile methods and available structural information. It is shown that all archaea encode a large number of proteins containing the helix-turn-helix (HTH) DNA-binding domains whose sequences are much more similar to bacterial HTH domains than to eukaryotic ones, such as the PAIRED, POU and homeodomains. The predominant class of HTH domains in archaea is the winged-HTH domain. The number and diversity of HTH domains in archaea is comparable to that seen in bacteria. The HTH domain in archaea combines with a variety of other domains that include replication system components, such as MCM proteins, translation system components, such as the alpha-subunit of phenyl-alanyl-tRNA synthetase, and several metabolic enzymes. The majority of the archaeal HTH-containing proteins are predicted to be gene/operon-specific transcriptional regulators. This apparent bacterial-type mode of transcription regulation is in sharp contrast to the eukaryote-like layout of the core transcription machinery in the archaea. In addition to the predicted bacterial-type transcriptional regulators, the HTH domain is conserved in archaeal and eukaryotic core transcription factors, such as TFIIB, TFIIE-alpha and MBF1. MBF1 is the only highly conserved, classical HTH domain that is vertically inherited in all archaea and eukaryotes. In contrast, while eukaryotic TFIIB and TFIIE-alpha possess forms of the HTH domain that are divergent in sequence, their archaeal counterparts contain typical HTH domains. It is shown that, besides the HTH domain, archaea encode unexpectedly large numbers of two other predicted DNA-binding domains, namely the Arc/MetJ domain and the Zn-ribbon. The core transcription regulators in archaea and eukaryotes (TFIIB/TFB, TFIIE-alpha and MBF1) and in bacteria (the sigma factors) share no similarity beyond the presence of distinct HTH domains. Thus HTH domains might have been independently recruited for a role in transcription regulation in the bacterial and archaeal/eukaryotic lineages. During subsequent evolution, the similarity between archaeal and bacterial gene/operon transcriptional regulators might have been established and maintained through multiple horizontal gene transfer events.

Archaea↗

Genome sequence of the radioresistant bacterium Deinococcus radiodurans R1.

The complete genome sequence of the radiation-resistant bacterium Deinococcus radiodurans R1 is composed of two chromosomes (2,648,638 and 412,348 base pairs), a megaplasmid (177,466 base pairs), and a small plasmid (45,704 base pairs), yielding a total genome of 3,284, 156 base pairs. Multiple components distributed on the chromosomes and megaplasmid that contribute to the ability of D. radiodurans to survive under conditions of starvation, oxidative stress, and high amounts of DNA damage were identified. Deinococcus radiodurans represents an organism in which all systems for DNA repair, DNA damage export, desiccation and starvation recovery, and genetic redundancy are present in one cell.

Bacterial Proteins↗

Human and mouse homologs of Escherichia coli DinB (DNA polymerase IV), members of the UmuC/DinB superfamily.

To understand the mechanisms underlying mutagenesis in eukaryotes better, we have cloned mouse and human homologs of the Escherichia coli dinB gene. E. coli dinB encodes DNA polymerase IV and greatly increases spontaneous mutations when overexpressed. The mouse and human DinB1 amino acid sequences share significant identity with E. coli DinB, including distinct motifs implicated in catalysis, suggesting conservation of the polymerase function. These proteins are members of a large superfamily of DNA damage-bypass replication proteins, including the E. coli proteins UmuC and DinB and the Saccharomyces cerevisiae proteins Rev1 and Rad30. In a phylogenetic tree, the mouse and human DinB1 proteins specifically group with E. coli DinB, suggesting a mitochondrial origin for these genes. The human DINB1 gene is localized to chromosome 5q13 and is widely expressed.

Amino Acid Sequence↗

Did DNA replication evolve twice independently?

DNA replication is central to all extant cellular organisms. There are substantial functional similarities between the bacterial and the archaeal/eukaryotic replication machineries, including but not limited to defined origins, replication bidirectionality, RNA primers and leading and lagging strand synthesis. However, several core components of the bacterial replication machinery are unrelated or only distantly related to the functionally equivalent components of the archaeal/eukaryotic replication apparatus. This is in sharp contrast to the principal proteins involved in transcription and translation, which are highly conserved in all divisions of life. We performed detailed sequence comparisons of the proteins that fulfill indispensable functions in DNA replication and classified them into four main categories with respect to the conservation in bacteria and archaea/eukaryotes: (i) non-homologous, such as replicative polymerases and primases; (ii) containing homologous domains but apparently non-orthologous and conceivably independently recruited to function in replication, such as the principal replicative helicases or proofreading exonucleases; (iii) apparently orthologous but poorly conserved, such as the sliding clamp proteins or DNA ligases; (iv) orthologous and highly conserved, such as clamp-loader ATPases or 5'-->3' exonucleases (FLAP nucleases). The universal conservation of some components of the DNA replication machinery and enzymes for DNA precursor biosynthesis but not the principal DNA polymerases suggests that the last common ancestor (LCA) of all modern cellular life forms possessed DNA but did not replicate it the way extant cells do. We propose that the LCA had a genetic system that contained both RNA and DNA, with the latter being produced by reverse transcription. Consequently, the modern-type system for double-stranded DNA replication likely evolved independently in the bacterial and archaeal/eukaryotic lineages.

Animals↗

The cytoplasmic helical linker domain of receptor histidine kinase and methyl-accepting proteins is common to many prokaryotic signalling proteins.

Mutations in the cytoplasmic linker regions of receptor histidine kinase and chemoreceptor proteins have been shown previously to significantly impair receptor functions. Here we demonstrate significant sequence similarities between these regions in numerous histidine kinases, methyl-accepting proteins, adenylyl cyclases and other prokaryotic signalling proteins. It is suggested that these 'HAMP domains' possess roles of regulating the phosphorylation or methylation of homodimeric receptors by transmitting the conformational changes in periplasmic ligand-binding domains to cytoplasmic signalling kinase and methyl-acceptor domains.

Adenylyl Cyclases↗

Eukaryotic signalling domain homologues in archaea and bacteria. Ancient ancestry and horizontal gene transfer.

Phyletic distributions of eukaryotic signalling domains were studied using recently developed sensitive methods for protein sequence analysis, with an emphasis on the detection and accurate enumeration of homologues in bacteria and archaea. A major difference was found between the distributions of enzyme families that are typically found in all three divisions of cellular life and non-enzymatic domain families that are usually eukaryote-specific. Previously undetected bacterial homologues were identified for# plant pathogenesis-related proteins, Pad1, von Willebrand factor type A, src homology 3 and YWTD repeat-containing domains. Comparisons of the domain distributions in eukaryotes and prokaryotes enabled distinctions to be made between the domains originating prior to the last common ancestor of all known life forms and those apparently originating as consequences of horizontal gene transfer events. A number of transfers of signalling domains from eukaryotes to bacteria were confidently identified, in contrast to only a single case of apparent transfer from eukaryotes to archaea.

Amino Acid Sequence↗

Gleaning non-trivial structural, functional and evolutionary information about proteins by iterative database searches.

Using a number of diverse protein families as test cases, we investigate the ability of the recently developed iterative sequence database search method, PSI-BLAST, to identify subtle relationships between proteins that originally have been deemed detectable only at the level of structure-structure comparison. We show that PSI-BLAST can detect many, though not all, of such relationships, but the success critically depends on the optimal choice of the query sequence used to initiate the search. Generally, there is a correlation between the diversity of the sequences detected in the first pass of database screening and the ability of a given query to detect subtle relationships in subsequent iterations. Accordingly, a thorough analysis of protein superfamilies at the sequence level is necessary in order to maximize the chances of gleaning non-trivial structural and functional inferences, as opposed to a single search, initiated, for example, with the sequence of a protein whose structure is available. This strategy is illustrated by several findings, each of which involves an unexpected structural prediction: (i) a number of previously undetected proteins with the HSP70-actin fold are identified, including a highly conserved and nearly ubiquitous family of metal-dependent proteases (typified by bacterial O-sialoglycoprotease) that represent an adaptation of this fold to a new type of enzymatic activity; (ii) we show that, contrary to the previous conclusions, ATP-dependent and NAD-dependent DNA ligases are confidently predicted to possess the same fold; (iii) the C-terminal domain of 3-phosphoglycerate dehydrogenase, which binds serine and is involved in allosteric regulation of the enzyme activity, is shown to typify a new superfamily of ligand-binding, regulatory domains found primarily in enzymes and regulators of amino acid and purine metabolism; (iv) the immunoglobulin-like DNA-binding domain previously identified in the structures of transcription factors NFkappaB and NFAT is shown to be a member of a distinct superfamily of intracellular and extracellular domains with the immunoglobulin fold; and (v) the Rag-2 subunit of the V-D-J recombinase is shown to contain a kelch-type beta-propeller domain which rules out its evolutionary relationship with bacterial transposases.

Actins↗

DNA polymerase beta-like nucleotidyltransferase superfamily: identification of three new families, classification and evolutionary history.

A detailed analysis of the polbeta superfamily of nucleotidyltransferases was performed using computer methods for iterative database search, multiple alignment, motif analysis and structural modeling. Three previously uncharacterized families of predicted nucleotidyltransferases are described. One of these new families includes small proteins found in all archaea and some bacteria that appear to consist of the minimal nucleotidyltransferase domain and may resemble the ancestral state of this superfamily. Another new family that is specifically related to eukaryotic polyA polymerases is typified by yeast Trf4p and Trf5p proteins that are involved in chromatin remodeling. The TRF family is represented by multiple members in all eukaryotes and may be involved in yet unknown nucleotide polymerization reactions required for maintenance of chromatin structure. Another new family of bacterial and archaeal nucleotidyltransferases is predicted to function in signal transduction since, in addition to the nucleotidyltransferase domain, these proteins contain ligand-binding domains. It is further shown that the catalytic domain of gamma proteobacterial adenylyl cyclases is homologous to the polbeta superfamily nucleotidyltransferases which emphasizes the general trend for the origin of signal-transducing enzymes from those involved in replication, repair and RNA processing. Classification of the polbeta superfamily into distinct families and examination of their phyletic distribution suggests that the evolution of this type of nucleotidyltransferases may have included bursts of rapid divergence linked to the emergence of new functions as well as a number of horizontal gene transfer events.

Adenylyl Cyclases↗

Conserved domains in DNA repair proteins and evolution of repair systems.

A detailed analysis of protein domains involved in DNA repair was performed by comparing the sequences of the repair proteins from two well-studied model organisms, the bacterium Escherichia coli and yeast Saccharomyces cerevisiae, to the entire sets of protein sequences encoded in completely sequenced genomes of bacteria, archaea and eukaryotes. Previously uncharacterized conserved domains involved in repair were identified, namely four families of nucleases and a family of eukaryotic repair proteins related to the proliferating cell nuclear antigen. In addition, a number of previously undetected occurrences of known conserved domains were detected; for example, a modified helix-hairpin-helix nucleic acid-binding domain in archaeal and eukaryotic RecA homologs. There is a limited repertoire of conserved domains, primarily ATPases and nucleases, nucleic acid-binding domains and adaptor (protein-protein interaction) domains that comprise the repair machinery in all cells, but very few of the repair proteins are represented by orthologs with conserved domain architecture across the three superkingdoms of life. Both the external environment of an organism and the internal environment of the cell, such as the chromatin superstructure in eukaryotes, seem to have a profound effect on the layout of the repair systems. Another factor that apparently has made a major contribution to the composition of the repair machinery is horizontal gene transfer, particularly the invasion of eukaryotic genomes by organellar genes, but also a number of likely transfer events between bacteria and archaea. Several additional general trends in the evolution of repair proteins were noticed; in particular, multiple, independent fusions of helicase and nuclease domains, and independent inactivation of enzymatic domains that apparently retain adaptor or regulatory functions.

Amino Acid Sequence↗