Search PubMed⌕ Search

Biomedical subjects

L Aravind

Publications and source records attributed to L Aravind.

At least 91 records · Page 5Linked to original sources

Extensive domain shuffling in transcription regulators of DNA viruses and implications for the origin of fungal APSES transcription factors.

BACKGROUND: Viral DNA-binding proteins have served as good models to study the biochemistry of transcription regulation and chromatin dynamics. Computational analysis of viral DNA-binding regulatory proteins and identification of their previously undetected homologs encoded by cellular genomes might lead to a better understanding of their function and evolution in both viral and cellular systems. RESULTS: The phyletic range and the conserved DNA-binding domains of the viral regulatory proteins of the poxvirus D6R/N1R and baculoviral Bro protein families have not been previously defined. Using computational analysis, we show that the amino-terminal module of the D6R/N1R proteins defines a novel, conserved DNA-binding domain (the KilA-N domain) that is found in a wide range of proteins of large bacterial and eukaryotic DNA viruses. The KilA-N domain is suggested to be homologous to the fungal DNA-binding APSES domain. We provide evidence for the KilA-N and APSES domains sharing a common fold with the nucleic acid-binding modules of the LAGLIDADG nucleases and the amino-terminal domains of the tRNA endonuclease. The amino-terminal module of the Bro proteins is another, distinct DNA-binding domain (the Bro-N domain) that is present in proteins whose domain architectures parallel those of the KilA-N domain-containing proteins. A detailed analysis of the KilA-N and Bro-N domains and the associated domains points to extensive domain shuffling and lineage-specific gene family expansion within DNA virus genomes. CONCLUSIONS: We define a large class of novel viral DNA-binding proteins and their cellular homologs and identify their domain architectures. On the basis of phyletic pattern analysis we present evidence for a probable viral origin of the fungus-specific cell-cycle regulatory transcription factors containing the APSES DNA-binding domain. We also demonstrate the extensive role of lineage-specific gene expansion and domain shuffling, within a limited set of approximately 24 domains, in the generation of the diversity of virus-specific regulatory proteins.

Amino Acid Sequence↗

MOSC domains: ancient, predicted sulfur-carrier domains, present in diverse metal-sulfur cluster biosynthesis proteins including Molybdenum cofactor sulfurases.

Using computational analysis, a novel superfamily of beta-strand-rich domains was identified in the Molybdenum cofactor sulfurase and several other proteins from both prokaryotes and eukaryotes. These MOSC domains contain an absolutely conserved cysteine and occur either as stand-alone forms such as the bacterial YiiM proteins, or fused to other domains such as a NifS-like catalytic domain in Molybdenum cofactor sulfurase. The MOSC domain is predicted to be a sulfur-carrier domain that receives sulfur abstracted by the pyridoxal phosphate-dependent NifS-like enzymes, on its conserved cysteine, and delivers it for the formation of diverse sulfur-metal clusters. The identification of this domain may clarify the mechanism of biogenesis of various metallo-enzymes including Molybdenum cofactor-containing enzymes that are compromised in human type II xanthinuria.

Amino Acid Sequence↗

A DNA repair system specific for thermophilic Archaea and bacteria predicted by genomic context analysis.

During a systematic analysis of conserved gene context in prokaryotic genomes, a previously undetected, complex, partially conserved neighborhood consisting of more than 20 genes was discovered in most Archaea (with the exception of Thermoplasma acidophilum and Halobacterium NRC-1) and some bacteria, including the hyperthermophiles Thermotoga maritima and Aquifex aeolicus. The gene composition and gene order in this neighborhood vary greatly between species, but all versions have a stable, conserved core that consists of five genes. One of the core genes encodes a predicted DNA helicase, often fused to a predicted HD-superfamily hydrolase, and another encodes a RecB family exonuclease; three core genes remain uncharacterized, but one of these might encode a nuclease of a new family. Two more genes that belong to this neighborhood and are present in most of the genomes in which the neighborhood was detected encode, respectively, a predicted HD-superfamily hydrolase (possibly a nuclease) of a distinct family and a predicted, novel DNA polymerase. Another characteristic feature of this neighborhood is the expansion of a superfamily of paralogous, uncharacterized proteins, which are encoded by at least 20-30% of the genes in the neighborhood. The functional features of the proteins encoded in this neighborhood suggest that they comprise a previously undetected DNA repair system, which, to our knowledge, is the first repair system largely specific for thermophiles to be identified. This hypothetical repair system might be functionally analogous to the bacterial-eukaryotic system of translesion, mutagenic repair whose central components are DNA polymerases of the UmuC-DinB-Rad30-Rev1 superfamily, which typically are missing in thermophiles.

Amino Acid Sequence↗

Trends in protein evolution inferred from sequence and structure analysis.

Complementary developments in comparative genomics, protein structure determination and in-depth comparison of protein sequences and structures have provided a better understanding of the prevailing trends in the emergence and diversification of protein domains. The investigation of deep relationships among different classes of proteins involved in key cellular functions, such as nucleic acid polymerases and other nucleotide-dependent enzymes, indicates that a substantial set of diverse protein domains evolved within the primordial, ribozyme-dominated RNA world.

Evolution, Molecular↗

SWIM, a novel Zn-chelating domain present in bacteria, archaea and eukaryotes.

A previously undetected domain with a CxCx(n)CxH pattern of predicted zinc-chelating residues was identified in a variety of prokaryotic and eukaryotic proteins. These include bacterial ATPases of the SWI2/SNF2 family, plant MuDR transposases and transposase-derived Far1 nuclear proteins, and vertebrate MEK kinase-1. This domain was designated SWIM after SWI2/SNF2 and MuDR, and is predicted to have DNA-binding and protein-protein interaction functions in different contexts.

Amino Acid Sequence↗

Origin and evolution of eukaryotic apoptosis: the bacterial connection.

The availability of numerous complete genome sequences of prokaryotes and several eukaryotic genome sequences provides for new insights into the origin of unique functional systems of the eukaryotes. Several key enzymes of the apoptotic machinery, including the paracaspase and metacaspase families of the caspase-like protease superfamily, apoptotic ATPases and NACHT family NTPases, and mitochondrial HtrA-like proteases, have diverse homologs in bacteria, but not in archaea. Phylogenetic analysis strongly suggests a mitochondrial origin for metacaspases and the HtrA-like proteases, whereas acquisition from Actinomycetes appears to be the most likely scenario for AP-ATPases. The homologs of apoptotic proteins are particularly abundant and diverse in bacteria that undergo complex development, such as Actinomycetes, Cyanobacteria and alpha-proteobacteria, the latter being progenitors of the mitochondria. In these bacteria, the apoptosis-related domains typically form multidomain proteins, which are known or inferred to participate in signal transduction and regulation of gene expression. Some of these bacterial multidomain proteins contain fusions between apoptosis-related domains, such as AP-ATPase fused with a metacaspase or a TIR domain. Thus, bacterial homologs of eukaryotic apoptotic machinery components might functionally and physically interact with each other as parts of signaling pathways that remain to be investigated. An emerging scenario of the origin of the eukaryotic apoptotic system involves acquisition of several central apoptotic effectors as a consequence of mitochondrial endosymbiosis and probably also as a result of subsequent, additional horizontal gene transfer events, which was followed by recruitment of newly emerging eukaryotic domains as adaptors.

Animals↗

The role of lineage-specific gene family expansion in the evolution of eukaryotes.

A computational procedure was developed for systematic detection of lineage-specific expansions (LSEs) of protein families in sequenced genomes and applied to obtain a census of LSEs in five eukaryotic species, the yeasts Saccharomyces cerevisiae and Schizosaccharomyces pombe, the nematode Caenorhabditis elegans, the fruit fly Drosophila melanogaster, and the green plant Arabidopsis thaliana. A significant fraction of the proteins encoded in each of these genomes, up to 80% in A. thaliana, belong to LSEs. Many paralogous gene families in each of the analyzed species are almost entirely comprised of LSEs, indicating that their diversification occurred after the divergence of the major lineages of the eukaryotic crown group. The LSEs show readily discernible patterns of protein functions. The functional categories most prone to LSE are structural proteins, enzymes involved in an organism's response to pathogens and environmental stress, and various components of signaling pathways responsible for specificity, including ubiquitin ligase E3 subunits and transcription factors. The functions of several previously uncharacterized, vastly expanded protein families were predicted through in-depth protein sequence analysis, for example, small-molecule kinases and methylases that are expanded independently in the fly and in the nematode. The functions of several other major LSEs remain mysterious; these protein families are attractive targets for experimental discovery of novel, lineage-specific functions in eukaryotes. LSEs seem to be one of the principal means of adaptation and one of the most important sources of organizational and regulatory diversity in crown-group eukaryotes.

Animals↗

Comparative genomic analysis of archaeal genotypic variants in a single population and in two different oceanic provinces.

Planktonic crenarchaeotes are present in high abundance in Antarctic winter surface waters, and they also make up a large proportion of total cell numbers throughout deep ocean waters. To better characterize these uncultivated marine crenarchaeotes, we analyzed large genome fragments from individuals recovered from a single Antarctic picoplankton population and compared them to those from a representative obtained from deeper waters of the temperate North Pacific. Sequencing and analysis of the entire DNA insert from one Antarctic marine archaeon (fosmid 74A4) revealed differences in genome structure and content between Antarctic surface water and temperate deepwater archaea. Analysis of the predicted gene products encoded by the 74A4 sequence and those derived from a temperate, deepwater planktonic crenarchaeote (fosmid 4B7) revealed many typical archaeal proteins but also several proteins that so far have not been detected in archaea. The unique fraction of marine archaeal genes included, among others, those for a predicted RNA-binding protein of the bacterial cold shock family and a eukaryote-type Zn finger protein. Comparison of closely related archaea originating from a single population revealed significant genomic divergence that was not evident from 16S rRNA sequence variation. The data suggest that considerable functional diversity may exist within single populations of coexisting microbial strains, even those with identical 16S rRNA sequences. Our results also demonstrate that genomic approaches can provide high-resolution information relevant to microbial population genetics, ecology, and evolution, even for microbes that have not yet been cultivated.

Amino Acid Sequence↗

Quod erat demonstrandum? The mystery of experimental validation of apparently erroneous computational analyses of protein sequences.

BACKGROUND: Computational predictions are critical for directing the experimental study of protein functions. Therefore it is paradoxical when an apparently erroneous computational prediction seems to be supported by experiment. RESULTS: We analyzed six cases where application of novel or conventional computational methods for protein sequence and structure analysis led to non-trivial predictions that were subsequently supported by direct experiments. We show that, on all six occasions, the original prediction was unjustified, and in at least three cases, an alternative, well-supported computational prediction, incompatible with the original one, could be derived. The most unusual cases involved the identification of an archaeal cysteinyl-tRNA synthetase, a dihydropteroate synthase and a thymidylate synthase, for which experimental verifications of apparently erroneous computational predictions were reported. Using sequence-profile analysis, multiple alignment and secondary-structure prediction, we have identified the unique archaeal 'cysteinyl-tRNA synthetase' as a homolog of extracellular polygalactosaminidases, and the 'dihydropteroate synthase' as a member of the beta-lactamase-like superfamily of metal-dependent hydrolases. CONCLUSIONS: In each of the analyzed cases, the original computational predictions could be refuted and, in some instances, alternative strongly supported predictions were obtained. The nature of the experimental evidence that appears to support these predictions remains an open question. Some of these experiments might signify discovery of extremely unusual forms of the respective enzymes, whereas the results of others could be due to artifacts.

Acetyltransferases↗

Peptide-N-glycanases and DNA repair proteins, Xp-C/Rad4, are, respectively, active and inactivated enzymes sharing a common transglutaminase fold.

Yeast RAD4, its human ortholog Xp-C and their orthologs in other eukaryotes are DNA repair proteins which participate in nucleotide excision repair through a ubiquitin-dependent process. However, no conserved globular domains that might have shed light on their origin or functions have been reported for these proteins. By using sequence profile analysis, we show that RAD4/Xp-C proteins contain the ancient transglutaminase fold and are specifically related to the recently characterized peptide-N-glycanases (PNGases) which remove glycans from glycoproteins during their degradation. The PNGases retain the catalytic triad that is typical of this fold and are predicted to have a reaction mechanism similar to that involved in transglutamination. In contrast, the RAD4/Xp-C proteins are predicted to be inactive and are likely to only possess the protein interaction function in DNA repair. These proteins also contain a long, low-complexity insert in the globular transglutaminase domain. The RAD4/Xp-C proteins, along with other inactive transglutaminase-fold proteins, represent a case of functional re-assignment of an ancient domain following the loss of the ancestral enzymatic activity.

Amidohydrolases↗

Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements.

PSI-BLAST is an iterative program to search a database for proteins with distant similarity to a query sequence. We investigated over a dozen modifications to the methods used in PSI-BLAST, with the goal of improving accuracy in finding true positive matches. To evaluate performance we used a set of 103 queries for which the true positives in yeast had been annotated by human experts, and a popular measure of retrieval accuracy (ROC) that can be normalized to take on values between 0 (worst) and 1 (best). The modifications we consider novel improve the ROC score from 0.758 +/- 0.005 to 0.895 +/- 0.003. This does not include the benefits from four modifications we included in the 'baseline' version, even though they were not implemented in PSI-BLAST version 2.0. The improvement in accuracy was confirmed on a small second test set. This test involved analyzing three protein families with curated lists of true positives from the non-redundant protein database. The modification that accounts for the majority of the improvement is the use, for each database sequence, of a position-specific scoring system tuned to that sequence's amino acid composition. The use of composition-based statistics is particularly beneficial for large-scale automated applications of PSI-BLAST.

Algorithms↗

Adaptations of the helix-grip fold for ligand binding and catalysis in the START domain superfamily.

With a protein structure comparison, an iterative database search with sequence profiles, and a multiple-alignment analysis, we show that two domains with the helix-grip fold, the star-related lipid-transfer (START) domain of the MLN64 protein and the birch allergen, are homologous. They define a large, previously underappreciated superfamily that we call the START superfamily. In addition to the classical START domains that are primarily involved in eukaryotic signaling mediated by lipid binding and the birch antigen family that consists of plant proteins implicated in stress/pathogen response, the START superfamily includes bacterial polyketide cyclases/aromatases (e.g., TcmN and WhiE VI) and two families of previously uncharacterized proteins. The identification of this domain provides a structural prediction of an important class of enzymes involved in polyketide antibiotic synthesis and allows the prediction of their active site. It is predicted that all START domains contain a similar ligand-binding pocket. Modifications of this pocket determine the ligand-binding specificity and may also be the basis for at least two distinct enzymatic activities, those of a cyclase/aromatase and an RNase. Thus, the START domain superfamily is a rare case of the adaptation of a protein fold with a conserved ligand-binding mode for both a broad variety of catalytic activities and noncatalytic regulatory functions. Proteins 2001;43:134-144.

Allergens↗

Cloning the human and mouse MMS19 genes and functional complementation of a yeast mms19 deletion mutant.

The MMS19 gene of the yeast Saccharomyces cerevisiae encodes a polypeptide of unknown function which is required for both nucleotide excision repair (NER) and RNA polymerase II (RNAP II) transcription. Here we report the molecular cloning of human and mouse orthologs of the yeast MMS19 gene. Both human and Drosophila MMS19 cDNAs correct thermosensitive growth and sensitivity to killing by UV radiation in a yeast mutant deleted for the MMS19 gene, indicating functional conservation between the yeast and mammalian gene products. Alignment of the translated sequences of MMS19 from multiple eukaryotes, including mouse and human, revealed the presence of several conserved regions, including a HEAT repeat domain near the C-terminus. The presence of HEAT repeats, coupled with functional complementation of yeast mutant phenotypes by the orthologous protein from higher eukaryotes, suggests a role of Mms19 protein in the assembly of a multiprotein complex(es) required for NER and RNAP II transcription. Both the mouse and human genes are ubiquitously expressed as multiple transcripts, some of which appear to derive from alternative splicing. The ratio of different transcripts varies in several different tissue types.

Alternative Splicing↗

Regulatory potential, phyletic distribution and evolution of ancient, intracellular small-molecule-binding domains.

Central cellular functions such as metabolism, solute transport and signal transduction are regulated, in part, via binding of small molecules by specialized domains. Using sensitive methods for sequence profile analysis and protein structure comparison, we exhaustively surveyed the protein sets from completely sequenced genomes for all occurrences of 21 intracellular small-molecule-binding domains (SMBDs) that are represented in at least two of the three major divisions of life (bacteria, archaea and eukaryotes). These included previously characterized domains such as PAS, GAF, ACT and ferredoxins, as well as three newly predicted SMBDs, namely the 4-vinyl reductase (4VR) domain, the NIFX domain and the 3-histidines (3H) domain. Although there are only a limited number of different superfamilies of these ancient SMBDs, they are present in numerous distinct proteins combined with various enzymatic, transport and signal-transducing domains. Most of the SMBDs show considerable evolutionary mobility and are involved in the generation of many lineage-specific domain architectures. Frequent re-invention of analogous architectures involving functionally related, but not homologous, domains was detected, such as, fusion of different SMBDs to several types of DNA-binding domains to form diverse transcription regulators in prokaryotes and eukaryotes. This is suggestive of similar selective forces affecting the diverse SMBDs and resulting in the formation of multidomain proteins that fit a limited number of functional stereotypes. Using the "guilt by association approach", the identification of SMBDs allowed prediction of functions and mode of regulation for a variety of previously uncharacterized proteins.

Amino Acid Sequence↗

TRAM, a predicted RNA-binding domain, common to tRNA uracil methylation and adenine thiolation enzymes.

A previously undetected conserved domain is identified in two distinct classes of tRNA-modifying enzymes, namely uridine methylases of the TRM2 family and enzymes of the MiaB family that are involved in 2-methylthioadenine formation. This domain, for which the acronym TRAM is proposed after TRM2 and MiaB, is predicted to bind tRNA and deliver the RNA-modifying enzymatic domains to their targets. In addition to the two families of RNA-modifying enzymes, the TRAM domain is present in several other proteins associated with the translation machinery and in a family of small, uncharacterized archaeal proteins that are predicted to have a role in the regulation of tRNA modification or translation. Secondary structure prediction indicates that the TRAM domain adopts a simple beta-barrel fold. In addition, sequence analysis of the MiaB family enzymes showed that they share the predicted catalytic site with biotin and lipoate synthases and probably employ the same mechanism for sulfur insertion into their respective substrate.

Adenine↗

The DNA-repair protein AlkB, EGL-9, and leprecan define new families of 2-oxoglutarate- and iron-dependent dioxygenases.

BACKGROUND: Protein fold recognition using sequence profile searches frequently allows prediction of the structure and biochemical mechanisms of proteins with an important biological function but unknown biochemical activity. Here we describe such predictions resulting from an analysis of the 2-oxoglutarate (2OG) and Fe(II)-dependent oxygenases, a class of enzymes that are widespread in eukaryotes and bacteria and catalyze a variety of reactions typically involving the oxidation of an organic substrate using a dioxygen molecule. RESULTS: We employ sequence profile analysis to show that the DNA repair protein AlkB, the extracellular matrix protein leprecan, the disease-resistance-related protein EGL-9 and several uncharacterized proteins define novel families of enzymes of the 2OG-Fe(II) oxygenase superfamily. The identification of AlkB as a member of the 2OG-Fe(II) oxygenase superfamily suggests that this protein catalyzes oxidative detoxification of alkylated bases. More distant homologs of AlkB were detected in eukaryotes and in plant RNA viruses, leading to the hypothesis that these proteins might be involved in RNA demethylation. The EGL-9 protein from Caenorhabditis elegans is necessary for normal muscle function and its inactivation results in resistance against paralysis induced by the Pseudomonas aeruginosa toxin. EGL-9 and leprecan are predicted to be novel protein hydroxylases that might be involved in the generation of substrates for protein glycosylation. CONCLUSIONS: Here, using sequence profile searches, we show that several previously undetected protein families contain 2OG-Fe(II) oxygenase fold. This allows us to predict the catalytic activity for a wide range of biologically important, but biochemically uncharacterized proteins from eukaryotes and bacteria.

Amino Acid Sequence↗