Search PubMed⌕ Search

Biomedical subjects

L Aravind

Publications and source records attributed to L Aravind.

At least 163 records · Page 9Linked to original sources

Fold prediction and evolutionary analysis of the POZ domain: structural and evolutionary relationship with the potassium channel tetramerization domain.

Using iterative database searches, a statistically significant sequence similarity was detected between the POZ (poxvirus and zinc finger) domains found in a variety of proteins involved in animal transcription regulation, cytoskeleton organization, and development, and the tetramerization domain of animal potassium channels. Using the crystal structure of the Aplysia Shaker channel tetramerization domain as a template, the common structure of the POZ domain class was predicted. Examination of the structure resulted in the identification of several structural features and specific amino acid residues that may be involved in conserved protein-protein interactions mediated by the POZ domains as well as those that may contribute to the specificity of these interactions. Phylogenetic analysis of the POZ domains suggests that the common ancestor of the crown group eukaryotes already possessed this domain; POZ domains have undergone independent expansion in plants and in different animal lineages.

Amino Acid Sequence↗

Novel predicted RNA-binding domains associated with the translation machinery.

Two previously undetected domains were identified in a variety of RNA-binding proteins, particularly RNA-modifying enzymes, using methods for sequence profile analysis. A small domain consisting of 60-65 amino acid residues was detected in the ribosomal protein S4, two families of pseudouridine synthases, a novel family of predicted RNA methylases, a yeast protein containing a pseudouridine synthetase and a deaminase domain, bacterial tyrosyl-tRNA synthetases, and a number of uncharacterized, small proteins that may be involved in translation regulation. Another novel domain, designated PUA domain, after PseudoUridine synthase and Archaeosine transglycosylase, was detected in archaeal and eukaryotic pseudouridine synthases, archaeal archaeosine synthases, a family of predicted ATPases that may be involved in RNA modification, a family of predicted archaeal and bacterial rRNA methylases. Additionally, the PUA domain was detected in a family of eukaryotic proteins that also contain a domain homologous to the translation initiation factor eIF1/SUI1; these proteins may comprise a novel type of translation factors. Unexpectedly, the PUA domain was detected also in bacterial and yeast glutamate kinases; this is compatible with the demonstrated role of these enzymes in the regulation of the expression of other genes. We propose that the S4 domain and the PUA domain bind RNA molecules with complex folded structures, adding to the growing collection of nucleic acid-binding domains associated with DNA and RNA modification enzymes. The evolution of the translation machinery components containing the S4, PUA, and SUI1 domains must have included several events of lateral gene transfer and gene loss as well as lineage-specific domain fusions.

Amino Acid Sequence↗

Reply

Explore the source record for details and available documents.

Journal Article↗

Origin of multicellular eukaryotes - insights from proteome comparisons.

The complete genomes of the yeast Saccharomyces cerevisiae and the nematode worm Caenorhabditis elegans have recently become available allowing the comparison of the complete protein sets of a unicellular and multicellular eukaryote for the first time. These comparisons reveal some striking trends in terms of expansions or extensive shuffling of specific domains that are involved in regulatory functions and signaling. Similar comparisons with the available sequence data from the plant Arabidopsis thaliana produce consistent results. These observations have provided useful insights regarding the origin of multicellular organisms.

Animals↗

The domains of death: evolution of the apoptosis machinery.

Recent progress in research into programmed cell death has resulted in the identification of the principal protein domains involved in this process. The evolution of many of these domains can be traced back in evolution to unicellular eukaryotes or even bacteria, where the domains appear to be involved in other regulatory functions. Cell-death systems in animals and plants share several conserved domains, in particular the family of apoptotic ATPases; this allows us to suggest a plausible, even if still incomplete, scenario for the evolution of apoptosis.

Adenosine Triphosphatases↗

IMPALA: matching a protein sequence against a collection of PSI-BLAST-constructed position-specific score matrices.

MOTIVATION: Many studies have shown that database searches using position-specific score matrices (PSSMs) or profiles as queries are more effective at identifying distant protein relationships than are searches that use simple sequences as queries. One popular program for constructing a PSSM and comparing it with a database of sequences is Position-Specific Iterated BLAST (PSI-BLAST). RESULTS: This paper describes a new software package, IMPALA, designed for the complementary procedure of comparing a single query sequence with a database of PSI-BLAST-generated PSSMs. We illustrate the use of IMPALA to search a database of PSSMs for protein folds, and one for protein domains involved in signal transduction. IMPALA's sensitivity to distant biological relationships is very similar to that of PSI-BLAST. However, IMPALA employs a more refined analysis of statistical significance and, unlike PSI-BLAST, guarantees the output of the optimal local alignment by using the rigorous Smith-Waterman algorithm. Also, it is considerably faster when run with a large database of PSSMs than is BLAST or PSI-BLAST when run against the complete non-redundant protein database.

Algorithms↗

A superfamily of archaeal, bacterial, and eukaryotic proteins homologous to animal transglutaminases.

Computer analysis using profiles generated by the PSI-BLAST program identified a superfamily of proteins homologous to eukaryotic transglutaminases. The members of the new protein superfamily are found in all archaea, show a sporadic distribution among bacteria, and were detected also in eukaryotes, such as two yeast species and the nematode Caenorhabditis elegans. Sequence conservation in this superfamily primarily involves three motifs that center around conserved cysteine, histidine, and aspartate residues that form the catalytic triad in the structurally characterized transglutaminase, the human blood clotting factor XIIIa'. On the basis of the experimentally demonstrated activity of the Methanobacterium phage pseudomurein endoisopeptidase, it is proposed that many, if not all, microbial homologs of the transglutaminases are proteases and that the eukaryotic transglutaminases have evolved from an ancestral protease.

Amino Acid Sequence↗

Comparative genomics of the Archaea (Euryarchaeota): evolution of conserved protein families, the stable core, and the variable shell.

Comparative analysis of the protein sequences encoded in the four euryarchaeal species whose genomes have been sequenced completely (Methanococcus jannaschii, Methanobacterium thermoautotrophicum, Archaeoglobus fulgidus, and Pyrococcus horikoshii) revealed 1326 orthologous sets, of which 543 are represented in all four species. The proteins that belong to these conserved euryarchaeal families comprise 31%-35% of the gene complement and may be considered the evolutionarily stable core of the archaeal genomes. The core gene set includes the great majority of genes coding for proteins involved in genome replication and expression, but only a relatively small subset of metabolic functions. For many gene families that are conserved in all euryarchaea, previously undetected orthologs in bacteria and eukaryotes were identified. A number of euryarchaeal synapomorphies (unique shared characters) were identified; these are protein families that possess sequence signatures or domain architectures that are conserved in all euryarchaea but are not found in bacteria or eukaryotes. In addition, euryarchaea-specific expansions of several protein and domain families were detected. In terms of their apparent phylogenetic affinities, the archaeal protein families split into bacterial and eukaryotic families. The majority of the proteins that have only eukaryotic orthologs or show the greatest similarity to their eukaryotic counterparts belong to the core set. The families of euryarchaeal genes that are conserved in only two or three species constitute a relatively mobile component of the genomes whose evolution should have involved multiple events of lineage-specific gene loss and horizontal gene transfer. Frequently these proteins have detectable orthologs only in bacteria or show the greatest similarity to the bacterial homologs, which might suggest a significant role of horizontal gene transfer from bacteria in the evolution of the euryarchaeota.

Amino Acid Sequence↗

Evolution of aminoacyl-tRNA synthetases--analysis of unique domain architectures and phylogenetic trees reveals a complex history of horizontal gene transfer events.

Phylogenetic analysis of aminoacyl-tRNA synthetases (aaRSs) of all 20 specificities from completely sequenced bacterial, archaeal, and eukaryotic genomes reveals a complex evolutionary picture. Detailed examination of the domain architecture of aaRSs using sequence profile searches delineated a network of partially conserved domains that is even more elaborate than previously suspected. Several unexpected evolutionary connections were identified, including the apparent origin of the beta-subunit of bacterial GlyRS from the HD superfamily of hydrolases, a domain shared by bacterial AspRS and the B subunit of archaeal glutamyl-tRNA amidotransferases, and another previously undetected domain that is conserved in a subset of ThrRS, guanosine polyphosphate hydrolases and synthetases, and a family of GTPases. Comparison of domain architectures and multiple alignments resulted in the delineation of synapomorphies-shared derived characters, such as extra domains or inserts-for most of the aaRSs specificities. These synapomorphies partition sets of aaRSs with the same specificity into two or more distinct and apparently monophyletic groups. In conjunction with cluster analysis and a modification of the midpoint-rooting procedure, this partitioning was used to infer the likely root position in phylogenetic trees. The topologies of the resulting rooted trees for most of the aaRSs specificities are compatible with the evolutionary "standard model" whereby the earliest radiation event separated bacteria from the common ancestor of archaea and eukaryotes as opposed to the two other possible evolutionary scenarios for the three major divisions of life. For almost all aaRSs specificities, however, this simple scheme is confounded by displacement of some of the bacterial aaRSs by their eukaryotic or, less frequently, archaeal counterparts. Displacement of ancestral eukaryotic aaRS genes by bacterial ones, presumably of mitochondrial origin, was observed for three aaRSs. In contrast, there was no convincing evidence of displacement of archaeal aaRSs by bacterial ones. Displacement of aaRS genes by eukaryotic counterparts is most common among parasitic and symbiotic bacteria, particularly the spirochaetes, in which 10 of the 19 aaRSs seem to have been displaced by the respective eukaryotic genes and two by the archaeal counterpart. Unlike the primary radiation events between the three main divisions of life, that were readily traceable through the phylogenetic analysis of aaRSs, no consistent large-scale bacterial phylogeny could be established. In part, this may be due to additional gene displacement events among bacterial lineages. Argument is presented that, although lineage-specific gene loss might have contributed to the evolution of some of the aaRSs, this is not a viable alternative to horizontal gene transfer as the principal evolutionary phenomenon in this gene class.

Amino Acid Sequence↗

An evolutionary classification of the metallo-beta-lactamase fold proteins.

All the detectable metallo-beta-lactamase fold proteins were identified in the publicly available sequence databases and complete genome sequences using iterative profile searches with the PSI-BLAST program and motif searches with position specific weight matrices. The catalytic site/mechanism and the corresponding structural elements were characterized for these proteins based on the available structure of the Bacillus zinc-dependent beta-lactamase. Based on pair-wise sequence and phylogenetic analysis an evolutionary classification for enzymes of this fold was developed and discussed in terms of implications for substrate specificity. Finally, some predicted inactive members which have been recruited for non-enzymatic functions such as microtubule binding in a cytoskeletal MAP1 are described.

Amino Acid Sequence↗

AAA+: A class of chaperone-like ATPases associated with the assembly, operation, and disassembly of protein complexes.

Using a combination of computer methods for iterative database searches and multiple sequence alignment, we show that protein sequences related to the AAA family of ATPases are far more prevalent than reported previously. Among these are regulatory components of Lon and Clp proteases, proteins involved in DNA replication, recombination, and restriction (including subunits of the origin recognition complex, replication factor C proteins, MCM DNA-licensing factors and the bacterial DnaA, RuvB, and McrB proteins), prokaryotic NtrC-related transcription regulators, the Bacillus sporulation protein SpoVJ, Mg2+, and Co2+ chelatases, the Halobacterium GvpN gas vesicle synthesis protein, dynein motor proteins, TorsinA, and Rubisco activase. Alignment of these sequences, in light of the structures of the clamp loader delta' subunit of Escherichia coli DNA polymerase III and the hexamerization component of N-ethylmaleimide-sensitive fusion protein, provides structural and mechanistic insights into these proteins, collectively designated the AAA+ class. Whole-genome analysis indicates that this class is ancient and has undergone considerable functional divergence prior to the emergence of the major divisions of life. These proteins often perform chaperone-like functions that assist in the assembly, operation, or disassembly of protein complexes. The hexameric architecture often associated with this class can provide a hole through which DNA or RNA can be thread; this may be important for assembly or remodeling of DNA-protein complexes.

Adenosine Triphosphatases↗

Comparison of the complete protein sets of worm and yeast: orthology and divergence.

Comparative analysis of predicted protein sequences encoded by the genomes of Caenorhabditis elegans and Saccharomyces cerevisiae suggests that most of the core biological functions are carried out by orthologous proteins (proteins of different species that can be traced back to a common ancestor) that occur in comparable numbers. The specialized processes of signal transduction and regulatory control that are unique to the multicellular worm appear to use novel proteins, many of which re-use conserved domains. Major expansion of the number of some of these domains seen in the worm may have contributed to the advent of multicellularity. The proteins conserved in yeast and worm are likely to have orthologs throughout eukaryotes; in contrast, the proteins unique to the worm may well define metazoans.

Animals↗

Chromosome 2 sequence of the human malaria parasite Plasmodium falciparum.

Chromosome 2 of Plasmodium falciparum was sequenced; this sequence contains 947,103 base pairs and encodes 210 predicted genes. In comparison with the Saccharomyces cerevisiae genome, chromosome 2 has a lower gene density, introns are more frequent, and proteins are markedly enriched in nonglobular domains. A family of surface proteins, rifins, that may play a role in antigenic variation was identified. The complete sequencing of chromosome 2 has shown that sequencing of the A+T-rich P. falciparum genome is technically feasible.

Amino Acid Sequence↗