Search PubMed⌕ Search

Biomedical subjects

E V Koonin

Publications and source records attributed to E V Koonin.

At least 19 recordsLinked to original sources

SURVEY AND SUMMARY: holliday junction resolvases and related nucleases: identification of new families, phyletic distribution and evolutionary trajectories.

Holliday junction resolvases (HJRs) are key enzymes of DNA recombination. A detailed computer analysis of the structural and evolutionary relationships of HJRs and related nucleases suggests that the HJR function has evolved independently from at least four distinct structural folds, namely RNase H, endonuclease, endonuclease VII-colicin E and RusA. The endonuclease fold, whose structural prototypes are the phage lambda exonuclease, the very short patch repair nuclease (Vsr) and type II restriction enzymes, is shown to encompass by far a greater diversity of nucleases than previously suspected. This fold unifies archaeal HJRs, repair nucleases such as RecB and Vsr, restriction enzymes and a variety of predicted nucleases whose specific activities remain to be determined. Within the RNase H fold a new family of predicted HJRs, which is nearly ubiquitous in bacteria, was discovered, in addition to the previously characterized RuvC family. The proteins of this family, typified by Escherichia coli YqgF, are likely to function as an alternative to RuvC in most bacteria, but could be the principal HJRs in low-GC Gram-positive bacteria and AQUIFEX: Endonuclease VII of phage T4 is shown to serve as a structural template for many nucleases, including MCR:A and other type II restriction enzymes. Together with colicin E7, endonuclease VII defines a distinct metal-dependent nuclease fold. As a result of this analysis, the principal HJRs are now known or confidently predicted for all bacteria and archaea whose genomes have been completely sequenced, with many species encoding multiple potential HJRs. Horizontal gene transfer, lineage-specific gene loss and gene family expansion, and non-orthologous gene displacement seem to have been major forces in the evolution of HJRs and related nucleases. A remarkable case of displacement is seen in the Lyme disease spirochete Borrelia burgdorferi, which does not possess any of the typical HJRs, but instead encodes, in its chromosome and each of the linear plasmids, members of the lambda exonuclease family predicted to function as HJRs. The diversity of HJRs and related nucleases in bacteria and archaea contrasts with their near absence in eukaryotes. The few detected eukaryotic representatives of the endonuclease fold and the RNase H fold have probably been acquired from bacteria via horizontal gene transfer. The identity of the principal HJR(s) involved in recombination in eukaryotes remains uncertain; this function could be performed by topoisomerase IB or by a novel, so far undetected, class of enzymes. Likely HJRs and related nucleases were identified in the genomes of numerous bacterial and eukaryotic DNA viruses. Gene flow between viral and cellular genomes has probably played a major role in the evolution of this class of enzymes. This analysis resulted in the prediction of numerous previously unnoticed nucleases, some of which are likely to be new restriction enzymes.

Amino Acid Sequence↗

Bacterial rhodopsin: evidence for a new type of phototrophy in the sea.

Extremely halophilic archaea contain retinal-binding integral membrane proteins called bacteriorhodopsins that function as light-driven proton pumps. So far, bacteriorhodopsins capable of generating a chemiosmotic membrane potential in response to light have been demonstrated only in halophilic archaea. We describe here a type of rhodopsin derived from bacteria that was discovered through genomic analyses of naturally occuring marine bacterioplankton. The bacterial rhodopsin was encoded in the genome of an uncultivated gamma-proteobacterium and shared highest amino acid sequence similarity with archaeal rhodopsins. The protein was functionally expressed in Escherichia coli and bound retinal to form an active, light-driven proton pump. The new rhodopsin exhibited a photochemical reaction cycle with intermediates and kinetics characteristic of archaeal proton-pumping rhodopsins. Our results demonstrate that archaeal-like rhodopsins are broadly distributed among different taxa, including members of the domain Bacteria. Our data also indicate that a previously unsuspected mode of bacterially mediated light-driven energy generation may commonly occur in oceanic surface waters worldwide.

Aerobiosis↗

Bacterial-type DNA holliday junction resolvases in eukaryotic viruses.

Homologous DNA recombination promotes genetic diversity and the maintenance of genome integrity, yet no enzymes with specificity for the Holliday junction (HJ)-a key DNA recombination intermediate-have been purified and characterized from metazoa or their viruses. Here we identify critical structural elements of RuvC, a bacterial HJ resolvase, in uncharacterized open reading frames from poxviruses and an iridovirus. The putative vaccinia virus resolvase was expressed as a recombinant protein, affinity purified, and shown to specifically bind and cleave a synthetic HJ to yield nicked duplex molecules. Mutation of either of two conserved acidic amino acids abrogated the catalytic activity of the A22R protein without affecting HJ binding. The presence of bacterial-type enzymes in metazoan viruses raises evolutionary questions.

Amino Acid Sequence↗

Estimating the number of protein folds and families from complete genome data.

Using the data on proteins encoded in complete genomes, combined with a rigorous theory of the sampling process, we estimate the total number of protein folds and families, as well as the number of folds and families in each genome. The total number of folds in globular, water- soluble proteins is estimated at about 1000, with structural information currently available for about one-third of the number. The sequenced genomes of unicellular organisms encode from approximately 25%, for the minimal genomes of the Mycoplasmas, to 70-80% for larger genomes, such as Escherichia coli and yeast, of the total number of folds. The number of protein families with significant sequence conservation was estimated to be between 4000 and 7000, with structures available for about 20% of these.

Conserved Sequence↗

Functional implications from crystal structures of the conserved Bacillus subtilis protein Maf with and without dUTP.

Three-dimensional structures of functionally uncharacterized proteins may furnish insight into their functions. The potential benefits of three-dimensional structural information regarding such proteins are particularly obvious when the corresponding genes are conserved during evolution, implying an important function, and no functional classification can be inferred from their sequences. The Bacillus subtilis Maf protein is representative of a family of proteins that has homologs in many of the completely sequenced genomes from archaea, prokaryotes, and eukaryotes, but whose function is unknown. As an aid in exploring function, we determined the crystal structure of this protein at a resolution of 1.85 A. The structure, in combination with multiple sequence alignment, reveals a putative active site. Phosphate ions present at this site and structural similarities between a portion of Maf and the anticodon-binding domains of several tRNA synthetases suggest that Maf may be a nucleic acid-binding protein. The crystal structure of a Maf-nucleoside triphosphate complex provides support for this hypothesis and hints at di- or oligonucleotides with either 5'- or 3'-terminal phosphate groups as ligands or substrates of Maf. A further clue comes from the observation that the structure of the Maf monomer bears similarity to that of the recently reported Methanococcus jannaschii Mj0226 protein. Just as for Maf, the structure of this predicted NTPase was determined as part of a structural genomics pilot project. The structural relation between Maf and Mj0226 was not apparent from sequence analysis approaches. These results emphasize the potential of structural genomics to reveal new unexpected connections between protein families previously considered unrelated.

Amino Acid Sequence↗

Aldolases of the DhnA family: a possible solution to the problem of pentose and hexose biosynthesis in archaea.

Sequence analysis of the recently identified class I aldolase of Escherichia coli (dhnA gene product) helped to identify its homologs in Chlamydia trachomatis, Chlamydiophyla pneumoniae and in each of the completely sequenced archaeal genomes. Iterative database searches revealed sequence similarities between the DhnA-family enzymes, deoxyribose phosphate aldolases and bacterial (class II) fructose bisphosphate aldolases and allowed prediction of similar three-dimensional structures (TIM-barrel fold) in all these enzymes. The Schiff base-forming lysyl residues of DhnA and deoxyribose phosphate aldolase are conserved in all members of the DhnA and deoxyribose phosphate aldolase families, indicating that these enzymes share common features with both class I and class II aldolases. The DhnA-family enzymes are predicted to possess an aldolase activity and to play a critical role in sugar biosynthesis in archaea.

Amino Acid Sequence↗

A family of ubiquitin-like proteins binds the ATPase domain of Hsp70-like Stch.

We have isolated two human ubiquitin-like (UbL) proteins that bind to a short peptide within the ATPase domain of the Hsp70-like Stch protein. Chap1 is a duplicated homologue of the yeast Dsk2 gene that is required for transit through the G2/M phase of the cell cycle and expression of the human full-length cDNA restored viability and suppressed the G2/M arrest phenotype of dsk2Delta rad23Delta Saccharomyces cerevisiae mutants. Chap2 is a homologue for Xenopus scythe which is an essential component of reaper-induced apoptosis in egg extracts. While the N-terminal UbL domains were not essential for Stch binding, Chap1/Dsk2 contains a Sti1-like repeat sequence that is required for binding to Stch and is also conserved in the Hsp70 binding proteins, Hip and p60/Sti1/Hop. These findings extend the association between Hsp70 members and genes encoding UbL sequences and suggest a broader role for the Hsp70-like ATPase family in regulating cell cycle and cell death events.

Adaptor Proteins, Signal Transducing↗

Prediction of transcription regulatory sites in Archaea by a comparative genomic approach.

Intragenomic and intergenomic comparisons of upstream nucleotide sequences of archaeal genes were performed with the goal of predicting transcription regulatory sites (operators) and identifying likely regulons. Learning sets for the detection of regulatory sites were constructed using the available experimental data on archaeal transcription regulation or by analogy with known bacterial regulons, and further analysis was performed using iterative profile searches. The information content of the candidate signals detected by this method is insufficient for reliable predictions to be made. Therefore, this approach has to be complemented by examination of evolutionary conservation in different archaeal genomes. This combined strategy resulted in the prediction of a conserved heat shock regulon in all euryarchaea, a nitrogen fixation regulon in the methanogens Methanococcus jannaschii and Methanobacterium thermoautotrophicum and an aromatic amino acid regulon in M.thermoautotrophicum. Unexpectedly, the heat shock regulatory site was detected not only for genes that encode known chaperone proteins but also for archaeal histone genes. This suggests a possible function for archaeal histones in stress-related changes in DNA condensation. In addition, comparative analysis of the genomes of three Pyrococcus species resulted in the prediction of their purine metabolism and transport regulon. The results demonstrate the feasibility of prediction of at least some transcription regulatory sites by comparing poorly characterized prokaryotic genomes, particularly when several closely related genome sequences are available.

Archaea↗

The COG database: a tool for genome-scale analysis of protein functions and evolution.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The database of Clusters of Orthologous Groups of proteins (COGs) is an attempt on a phylogenetic classification of the proteins encoded in 21 complete genomes of bacteria, archaea and eukaryotes (http://www. ncbi.nlm. nih.gov/COG). The COGs were constructed by applying the criterion of consistency of genome-specific best hits to the results of an exhaustive comparison of all protein sequences from these genomes. The database comprises 2091 COGs that include 56-83% of the gene products from each of the complete bacterial and archaeal genomes and approximately 35% of those from the yeast Saccharomyces cerevisiae genome. The COG database is accompanied by the COGNITOR program that is used to fit new proteins into the COGs and can be applied to functional and phylogenetic annotation of newly sequenced genomes.

Database Management Systems↗