Search PubMed⌕ Search

Biomedical subjects

Nick V Grishin

Publications and source records attributed to Nick V Grishin.

14 recordsLinked to original sources

Sequence and structure classification of kinases.

Kinases are a ubiquitous group of enzymes that catalyze the phosphoryl transfer reaction from a phosphate donor (usually ATP) to a receptor substrate. Although all kinases catalyze essentially the same phosphoryl transfer reaction, they display remarkable diversity in their substrate specificity, structure, and the pathways in which they participate. In order to learn the relationship between structural fold and functional specificities in kinases, we have done a comprehensive survey of all available kinase sequences (>17,000) and classified them into 30 distinct families based on sequence similarities. Of these families, 19, covering nearly 98% of all sequences, fall into seven general structural folds for which three-dimensional structures are known. These fold groups include some of the most widespread protein folds, such as Rossmann fold, ferredoxin fold, ribonuclease H fold, and TIM beta/alpha-barrel. On the basis of this classification system, we examined the shared substrate binding and catalytic mechanisms as well as variations of these mechanisms in the same fold groups. Cases of convergent evolution of identical kinase activities occurring in different folds are discussed.

Amino Acid Sequence↗

Expanding the nitrogen regulatory protein superfamily: Homology detection at below random sequence identity.

Nitrogen regulatory (PII) proteins are signal transduction molecules involved in controlling nitrogen metabolism in prokaryots. PII proteins integrate the signals of intracellular nitrogen and carbon status into the control of enzymes involved in nitrogen assimilation. Using elaborate sequence similarity detection schemes, we show that five clusters of orthologs (COGs) and several small divergent protein groups belong to the PII superfamily and predict their structure to be a (betaalphabeta)(2) ferredoxin-like fold. Proteins from the newly emerged PII superfamily are present in all major phylogenetic lineages. The PII homologs are quite diverse, with below random (as low as 1%) pairwise sequence identities between some members of distant groups. Despite this sequence diversity, evidence suggests that the different subfamilies retain the PII trimeric structure important for ligand-binding site formation and maintain a conservation of conservations at residue positions important for PII function. Because most of the orthologous groups within the PII superfamily are composed entirely of hypothetical proteins, our remote homology-based structure prediction provides the only information about them. Analogous to structural genomics efforts, such prediction gives clues to the biological roles of these proteins and allows us to hypothesize about locations of functional sites on model structures or rationalize about available experimental information. For instance, conserved residues in one of the families map in close proximity to each other on PII structure, allowing for a possible metal-binding site in the proteins coded by the locus known to affect sensitivity to divalent metal ions. Presented analysis pushes the limits of sequence similarity searches and exemplifies one of the extreme cases of reliable sequence-based structure prediction. In conjunction with structural genomics efforts to shed light on protein function, our strategies make it possible to detect homology between highly diverse sequences and are aimed at understanding the most remote evolutionary connections in the protein world.

Amino Acid Sequence↗

Crystal structure of Haemophilus influenzae NadR protein. A bifunctional enzyme endowed with NMN adenyltransferase and ribosylnicotinimide kinase activities.

Haemophilus influenzae NadR protein (hiNadR) has been shown to be a bifunctional enzyme possessing both NMN adenylytransferase (NMNAT; EC ) and ribosylnicotinamide kinase (RNK; EC ) activities. Its function is essential for the growth and survival of H. influenzae and thus may present a new highly specific anti-infectious drug target. We have solved the crystal structure of hiNadR complexed with NAD using the selenomethionine MAD phasing method. The structure reveals the presence of two distinct domains. The N-terminal domain that hosts the NMNAT activity is closely related to archaeal NMNAT, whereas the C-terminal domain, which has been experimentally demonstrated to possess ribosylnicotinamide kinase activity, is structurally similar to yeast thymidylate kinase and several other P-loop-containing kinases. There appears to be no cross-talk between the two active sites. The bound NAD at the active site of the NMNAT domain reveals several critical interactions between NAD and the protein. There is also a second non-active-site NAD molecule associated with the C-terminal RNK domain that adopts a highly folded conformation with the nicotinamide ring stacking over the adenine base. Whereas the RNK domain of the hiNadR structure presented here is the first structural characterization of a ribosylnicotinamide kinase from any organism, the NMNAT domain of hiNadR defines yet another member of the pyridine nucleotide adenylyltransferase family.

Amino Acid Sequence↗

C-terminal domain of gyrase A is predicted to have a beta-propeller structure.

Two different type II topoisomerases are known in bacteria. DNA gyrase (Gyr) introduces negative supercoils into DNA. Topoisomerase IV (Par) relaxes DNA supercoils. GyrA and ParC subunits of bacterial type II topoisomerases are involved in breakage and reunion of DNA. The spatial structure of the C-terminal fragment in GyrA/ParC is not available. We infer homology between the C-terminal domain of GyrA/ParC and a regulator of chromosome condensation (RCC1), a eukaryotic protein that functions as a guanine-nucleotide-exchange factor for the nuclear G protein Ran. This homology, complemented by detection of 6 sequence repeats with 4 predicted beta-strands each in GyrA/ParC sequences, allows us to predict that the GyrA/ParC C-terminal domain folds into a 6-bladed beta-propeller. The prediction rationalizes available experimental data and sheds light on the spatial properties of the largest topoisomerase domain that lacks structural information.

Amino Acid Sequence↗

A DNA repair system specific for thermophilic Archaea and bacteria predicted by genomic context analysis.

During a systematic analysis of conserved gene context in prokaryotic genomes, a previously undetected, complex, partially conserved neighborhood consisting of more than 20 genes was discovered in most Archaea (with the exception of Thermoplasma acidophilum and Halobacterium NRC-1) and some bacteria, including the hyperthermophiles Thermotoga maritima and Aquifex aeolicus. The gene composition and gene order in this neighborhood vary greatly between species, but all versions have a stable, conserved core that consists of five genes. One of the core genes encodes a predicted DNA helicase, often fused to a predicted HD-superfamily hydrolase, and another encodes a RecB family exonuclease; three core genes remain uncharacterized, but one of these might encode a nuclease of a new family. Two more genes that belong to this neighborhood and are present in most of the genomes in which the neighborhood was detected encode, respectively, a predicted HD-superfamily hydrolase (possibly a nuclease) of a distinct family and a predicted, novel DNA polymerase. Another characteristic feature of this neighborhood is the expansion of a superfamily of paralogous, uncharacterized proteins, which are encoded by at least 20-30% of the genes in the neighborhood. The functional features of the proteins encoded in this neighborhood suggest that they comprise a previously undetected DNA repair system, which, to our knowledge, is the first repair system largely specific for thermophiles to be identified. This hypothetical repair system might be functionally analogous to the bacterial-eukaryotic system of translesion, mutagenic repair whose central components are DNA polymerases of the UmuC-DinB-Rad30-Rev1 superfamily, which typically are missing in thermophiles.

Amino Acid Sequence↗

Structure of human nicotinamide/nicotinic acid mononucleotide adenylyltransferase. Basis for the dual substrate specificity and activation of the oncolytic agent tiazofurin.

Nicotinamide/nicotinate mononucleotide (NMN/ NaMN)adenylyltransferase (NMNAT) is an indispensable enzyme in the biosynthesis of NAD(+) and NADP(+). Human NMNAT displays unique dual substrate specificity toward both NMN and NaMN, thus flexible in participating in both de novo and salvage pathways of NAD synthesis. Human NMNAT also catalyzes the rate-limiting step of the metabolic conversion of the anticancer agent tiazofurin to its active form tiazofurin adenine dinucleotide (TAD). The tiazofurin resistance is mainly associated with the low NMNAT activity in the cell. We have solved the crystal structures of human NMNAT in complex with NAD, deamido-NAD, and a non-hydrolyzable TAD analogue beta-CH(2)-TAD. These complex structures delineate the broad substrate specificity of the enzyme toward both NMN and NaMN and reveal the structural mechanism for adenylation of tiazofurin nucleotide. The crystal structure of human NMNAT also shows that it forms a barrel-like hexamer with the predicted nuclear localization signal sequence located on the outside surface of the barrel, supporting its functional role of interacting with the nuclear transporting proteins. The results from the analytical ultracentrifugation studies are consistent with the formation of a hexamer in solution under certain conditions.

Amino Acid Sequence↗

Evolution of the regulators of G-protein signaling multigene family in mouse and human.

The regulators of G-protein signaling (RGS) proteins are important regulatory and structural components of G-protein coupled receptor complexes. RGS proteins are GTPase activating proteins (GAPs) of Gi-and Gq-class Galpha proteins, and thereby accelerate signaling kinetics and termination. Here, we mapped the chromosomal positions of all 21 Rgs genes in mouse, and determined human RGS gene structures using genomic sequence from partially assembled bacterial artificial chromosomes (BACs) and Celera fragments. In mice and humans, 18 of 21 RGS genes are either tandemly duplicated or tightly linked to genes encoding other components of G-protein signaling pathways, including Galpha, Ggamma, receptors (GPCR), and receptor kinases (GPRK). A phylogenetic tree revealed seven RGS gene subfamilies in the yeast and metazoan genomes that have been sequenced. We propose that similar systematic analyses of all multigene families from human and other mammalian genomes will help complete the assembly and annotation of the human genome sequence.

Animals↗

Genome trees and the tree of life.

Genome comparisons indicate that horizontal gene transfer and differential gene loss are major evolutionary phenomena that, at least in prokaryotes, involve a large fraction, if not the majority, of genes. The extent of these events casts doubt on the feasibility of constructing a 'Tree of Life', because the trees for different genes often tell different stories. However, alternative approaches to tree construction that attempt to determine tree topology on the basis of comparisons of complete gene sets seem to reveal a phylogenetic signal that supports the three-domain evolutionary scenario and suggests the possibility of delineation of previously undetected major clades of prokaryotes. If the validity of these whole-genome approaches to tree building is confirmed by analyses of numerous new genomes, which are currently being sequenced at an increasing rate, it would seem that the concept of a universal 'species' tree is still appropriate. However, this tree should be reinterpreted as a prevailing trend in the evolution of genome-scale gene sets rather than as a complete picture of evolution.

Animals↗

Evolution of protein structures and functions.

Within the ever-expanding repertoire of known protein sequences and structures, many examples of evolving three-dimensional structures are emerging that illustrate the plasticity and robustness of protein folds. The mechanisms by which protein folds change often include the fusion of duplicated domains, followed by divergence through mutation. Such changes reflect both the stability of protein folds and the requirements of protein function.

Amino Acid Motifs↗

Sec61beta--a component of the archaeal protein secretory system.

Sec61p/SecYEG complexes mediate protein translocation across membranes and are present in both eukaryotes and bacteria. Whereas homologues of Sec61alpha/SecY and Sec61gamma/SecE exist in archaea, identification of the third component (Sec61beta or SecG) has remained elusive. Using PSI-BLAST, the archaeal counterpart of Sec61beta has been detected. With the identification of the Sec61beta motif, functions for a universal family of archaeal proteins can be predicted and the archaeal translocon system can be definitively detected.

Amino Acid Motifs↗

Crystal structures of E. coli nicotinate mononucleotide adenylyltransferase and its complex with deamido-NAD.

Nicotinamide/Nicotinate mononucleotide (NMN/NaMN) adenylyltransferase is an indispensable enzyme in both de novo biosynthesis and salvage of NAD+ and NADP+. In prokaryotes, it is absolutely required for cell survival, thus representing an attractive target for the development of new broad-spectrum antibacteria inhibitors. The crystal structures of E. coli NaMN adenylyltransferase (NMNAT) and its complex with deamido-NAD (NaAD) revealed that ligand binding causes large conformational changes in several loop regions around the active site. The enzyme specifically recognizes the deamidated pyridine nucleotide through interactions between nicotinate carboxylate with several protein main chain amides and a positive helix dipole. Comparison of E. coli NMNAT with those from archaeal organisms revealed extensive differences in the active site architecture, enzyme-ligand interaction mode, and bound dinucleotide conformations. The bacterial NaMN adenylyltransferase structures described here provide a foundation for structure-based design of specific inhibitors that may have therapeutic potential.

Amino Acid Sequence↗

Euclidian space and grouping of biological objects.

MOTIVATION: Biological objects tend to cluster into discrete groups. Objects within a group typically possess similar properties. It is important to have fast and efficient tools for grouping objects that result in biologically meaningful clusters. Protein sequences reflect biological diversity and offer an extraordinary variety of objects for polishing clustering strategies. Grouping of sequences should reflect their evolutionary history and their functional properties. Visualization of relationships between sequences is of no less importance. Tree-building methods are typically used for such visualization. An alternative concept to visualization is a multidimensional sequence space. In this space, proteins are defined as points and distances between the points reflect the relationships between the proteins. Such a space can also be a basis for model-based clustering strategies that typically produce results correlating better with biological properties of proteins. RESULTS: We developed an approach to classification of biological objects that combines evolutionary measures of their similarity with a model-based clustering procedure. We apply the methodology to amino acid sequences. On the first step, given a multiple sequence alignment, we estimate evolutionary distances between proteins measured in expected numbers of amino acid substitutions per site. These distances are additive and are suitable for evolutionary tree reconstruction. On the second step, we find the best fit approximation of the evolutionary distances by Euclidian distances and thus represent each protein by a point in a multidimensional space. The Euclidian space may be projected in two or three dimensions and the projections can be used to visualize relationships between proteins. On the third step, we find a non-parametric estimate of the probability density of the points and cluster the points that belong to the same local maximum of this density in a group. The number of groups is controlled by a sigma-parameter that determines the shape of the density estimate and the number of maxima in it. The grouping procedure outperforms commonly used methods such as UPGMA and single linkage clustering.

Algorithms↗

Side-chain modeling with an optimized scoring function.

Modeling side-chain conformations on a fixed protein backbone has a wide application in structure prediction and molecular design. Each effort in this field requires decisions about a rotamer set, scoring function, and search strategy. We have developed a new and simple scoring function, which operates on side-chain rotamers and consists of the following energy terms: contact surface, volume overlap, backbone dependency, electrostatic interactions, and desolvation energy. The weights of these energy terms were optimized to achieve the minimal average root mean square (rms) deviation between the lowest energy rotamer and real side-chain conformation on a training set of high-resolution protein structures. In the course of optimization, for every residue, its side chain was replaced by varying rotamers, whereas conformations for all other residues were kept as they appeared in the crystal structure. We obtained prediction accuracy of 90.4% for chi(1), 78.3% for chi(1 + 2), and 1.18 A overall rms deviation. Furthermore, the derived scoring function combined with a Monte Carlo search algorithm was used to place all side chains onto a protein backbone simultaneously. The average prediction accuracy was 87.9% for chi(1), 73.2% for chi(1 + 2), and 1.34 A rms deviation for 30 protein structures. Our approach was compared with available side-chain construction methods and showed improvement over the best among them: 4.4% for chi(1), 4.7% for chi(1 + 2), and 0.21 A for rms deviation. We hypothesize that the scoring function instead of the search strategy is the main obstacle in side-chain modeling. Additionally, we show that a more detailed rotamer library is expected to increase chi(1 + 2) prediction accuracy but may have little effect on chi(1) prediction accuracy.

Algorithms↗

Breaking the singleton of germination protease.

Germination protease (GPR) plays an important role in the germination of spores of Bacillus and Clostridium species. A few very similar GPRs form a singleton group without significant sequence similarities to any other proteins. Their active site locations and catalytic mechanisms are unclear, despite the recent 3-D structure determination of Bacillus megaterium GPR. Using structural comparison and sequence analysis, we show that GPR is homologous to bacterial hydrogenase maturation protease (HybD). HybD's activity relies on the recognition and binding of metal ions in Ni-Fe hydrogenase, its substrate. Two highly conserved motifs are shared among GPRs, hydrogenase maturation proteases, and another group of hypothetical proteins. Conservation of two acidic residues in all these homologs indicates that metal binding is important for their function. Our analysis helps localize the active site of GPRs and provides insight into the catalytic mechanisms of a superfamily of putative metal-regulated proteases.

Amino Acid Sequence↗